Running a Multi-Agent Economic System on a 3B Parameter Small Model: Thousand Token Wood Practical Report

AI News Flash: ‘Thousand Token Wood’ is a multi-agent economic simulation system submitted to the Build Small Hackathon, using the Qwen2.5-3B small model to power five forest animal characters trading five types of goods for stone currency in a fictional market. The entire system is deployed on Modal with vLLM, with Gradio for the frontend…

2026-06-06 · 3 min · 539 words · Judy

Nemotron 3.5 Content Safety: Building Customizable Multimodal Guardrails for Global Enterprise AI

NVIDIA just released Nemotron 3.5 Content Safety—a multimodal safety classifier built for enterprise AI apps. Based on Google Gemma 3 4B with LoRA fine-tuning, it runs on just 8GB+ VRAM. Its biggest upgrade? Unified multimodal evaluation—processing user prompts, images, and assistant replies in a single pass.

2026-06-05 · 3 min · 521 words · Judy

EVA-Bench Data 2.0 Benchmark Released: Covering 3 Domains, 121 Tools, and 213 Test Scenarios

AI News: ServiceNow AI Research Team releases EVA-Bench Data 2.0, an enterprise-grade benchmark designed for Voice Agent, now significantly expanded from a single domain to three enterprise scenarios: Aviation Customer Service Management (CSM), Enterprise IT Service Management (ITSM), and Healthcare HR Service Delivery (HRSD). The three domains together cover 213…

2026-06-04 · 3 min · 550 words · Judy

Task-Seeded Synthetic QA Data Generation for Nemotron Pre-training

AI News Flash: NVIDIA developed a five-stage Task-Seeded Synthetic Data Generation (Task-Seeded SDG) process for the Nemotron series, selecting ~70 public tasks (~700 subtasks) from lm-eval-harness, divided into knowledge-intensive (39 tasks, ~3M samples) and reasoning-intensive (34 tasks, ~1.5M samples) seed categories…

2026-06-04 · 3 min · 486 words · Judy

Beyond Large Language Models: The Key to Enterprise AI at Scale is Agent Logic

AI News Flash: IBM Research study reveals that the key to enterprise AI scaling isn’t bigger LLMs—it’s ‘Agent Logic’: a guidance layer built from software primitives like knowledge graphs, static program analysis, and algorithm decomposition. This mechanism compresses LLM context space while reducing hallucination rates and token consumption, making model behavior more controllable and costs more predictable.

2026-06-01 · 1 min · 177 words · Judy

JetBrains Releases Mellum2: 12B Parameter Mixture-of-Experts Architecture Developer-Focused Model

AI News Flash: JetBrains released Mellum2 on June 1, 2026—a 12-billion parameter open-source model based on Mixture-of-Experts (MoE) architecture, but it only activates 2.5 billion active parameters per inference, making inference over twice as fast as models of equivalent scale, significantly reducing deployment costs, released under Apache 2.0 license. Mellum2 isn’t positioned as a replacement for frontier large models, but rather as a ‘focused model’ in multi-model collaboration systems, handling high-frequency lightweight tasks including prompt classification, tool selection, context compression and summarization for RAG pipelines, sub-agent planning validation, and code completion. The model processes only text and code modalities, deliberately excluding multimodal capabilities to keep the architecture lean—particularly suitable for enterprises deploying in private environments to handle internal code and confidential data. Across multiple benchmarks including code generation, reasoning, science, and math, Mellum2 achieves competitive performance among open-source models of similar scale. The technical report has also been published on arXiv (编号 2605.31268), and model weights are available for download on HuggingFace.

2026-06-01 · 3 min · 451 words · Judy

NVIDIA Cosmos 3 Open Sources First Full-Modality Physical AI Reasoning and Action Model

AI News Flash: NVIDIA releases Cosmos 3, an open full-modality World Foundation Model designed for Physical AI, featuring integrated image generation, physical reasoning, and action output in a single architecture, replacing the previous separate deployment of Cosmos Predict, Transfer, Reason, Policy…

2026-06-01 · 2 min · 295 words · Judy
Get our weekly AI digest:

AI engineering, trading systems, automation — curated weekly. No spam.