MosaicLeaks Study: Can AI Research Agents Really Keep Secrets?

AI News: MosaicLeaks is a new study on deep research AI agent privacy leaks, revealing a vulnerability called the ‘mosaic effect’: when agents simultaneously access local private files and external network tools, each seemingly harmless search query can accumulate to allow observers to piece together enterprise secrets. The study uses a medical institution as an example: to complete a multi-step question, the agent first queried cloud migration milestones…

2026-06-19 · 3 min · 498 words · Judy

Deploying Hugging Face Hub Models to Physical Robot Hardware via Strands Agents and LeRobot

AI News Flash: AWS open-sourced Strands Robots SDK (Apache 2.0 license), deeply integrated with Hugging Face LeRobot framework, aiming to bridge the complete workflow from robot demonstration data collection to physical hardware deployment. Previously, this path required five independent tools for recording demos, training models, simulation testing, hardware deployment, and multi-bot coordination, with no communication between tools.

2026-06-17 · 3 min · 533 words · Judy

AI Agent Chains Two Hugging Face Spaces to Auto-Generate a 3D Paris Gallery

AI news brief: a developer had an AI agent independently produce all the assets for a 3D showcase site of Paris landmarks, without manually opening any image generation tool or 3D reconstruction software. The agent completed the task by directly chaining two Hugging Face Spaces: first calling ideogram-ai/ideogram4 to turn each landmark into a clean, specimen-style image on a black background using text prompts…

2026-06-09 · 3 min · 574 words · Judy

Building the Pakistan Notice Helper: Using AI to Tackle Local Scam Alerts

AI News Brief: Pakistan Notice Helper is a small AI safety tool built specifically for Pakistan’s local scam message problem, created by a developer for the Backyard AI track of the ‘Build Small’ hackathon. Pakistani users constantly get suspicious messages disguised as banks, courier companies, tax authorities, telecoms, or government agencies — spotting the fake isn’t the hard part, knowing what to do next is.

2026-06-08 · 3 min · 596 words · Judy

Thousand Token Wood: Running a Multi-Agent Economy on a 3B Parameter Model

AI News Brief: Thousand Token Wood is a multi-agent economic simulation submitted to the Build Small Hackathon, using the small Qwen2.5-3B model to drive five forest animal characters trading five goods for pebble currency in a fictional market. The whole system runs on vLLM deployed on Modal, with a Gradio frontend…

2026-06-06 · 3 min · 582 words · Judy

Nemotron 3.5 Content Safety: Customizable Multimodal Safety for Global Enterprise AI

AI News Brief: NVIDIA has released Nemotron 3.5 Content Safety, a multimodal safety classifier built for enterprise AI applications, based on Google Gemma 3 4B with LoRA fine-tuning and deployable with just 8GB+ VRAM. The biggest difference from the previous generation is ‘unified multimodal evaluation’…

2026-06-05 · 3 min · 597 words · Judy

EVA-Bench Data 2.0 Benchmark Release: Covering 3 Domains, 121 Tools, and 213 Test Scenarios

AI News Flash: ServiceNow AI’s research team has released EVA-Bench Data 2.0, an enterprise-grade benchmark designed specifically for voice agents. This release dramatically expands the scope, moving from a single domain to three major enterprise scenarios: airline customer service management (CSM), enterprise IT service management (ITSM), and healthcare human resources service delivery (HRSD). Together, the three domains cover 2…

2026-06-04 · 3 min · 593 words · Judy

Task-Seeded Synthetic QA Data Generation for Nemotron Pretraining

AI News Brief: NVIDIA developed a five-stage ‘Task-Seeded SDG’ pipeline for the Nemotron model family — pulling roughly 70 public tasks (about 700 subtasks) from lm-eval-harness, split into knowledge-intensive (39 tasks, ~3M samples) and reasoning-intensive (34 tasks, ~1.5M samples) seed categories…

2026-06-04 · 3 min · 521 words · Judy

Beyond LLMs: Agent Logic Is the Real Key to Scaling Enterprise AI

AI News Flash: IBM Research published a study arguing that the key to scaling enterprise AI isn’t a bigger LLM — it’s ‘Agent Logic,’ a guidance layer built from software primitives like knowledge graphs, static code analysis, and algorithmic decomposition. This mechanism compresses the LLM’s context space, cutting both hallucination rates and token consumption while making model behavior more controllable and costs more predictable.

2026-06-01 · 3 min · 541 words · Judy

JetBrains Releases Mellum2: 12B-Parameter Mixture-of-Experts Model Built for Developers

AI news flash: On June 1, 2026, JetBrains released Mellum2, an open-source 12-billion-parameter model built on a mixture-of-experts (MoE) architecture that only activates 2.5 billion parameters per inference, making it more than twice as fast as models of similar size while cutting deployment costs significantly — released under the Apache 2.0 license.

2026-06-01 · 3 min · 440 words · Judy
Get our weekly AI digest:

AI engineering, trading systems, automation — curated weekly. No spam.