NVIDIA fits 512 experts in quarter the space


Quick Scribbles

  • NVIDIA — LatentMoE fits 512 experts in 4× less space, cutting memory costs 75%.
  • Google AX — Egress port rules aren't enforced at runtime, creating security gaps.
  • TypeSafe Jev — New model returns typed decisions 100× faster than LLM text generation.
  • Hugging Face — Hired MLX creator Jun Kim to accelerate Apple Silicon AI development.

Stay Connected

Subscribe to BrainScriblr for the latest AI developments delivered to your inbox.


Good morning, AI Knowledge Worker. NVIDIA published LatentMoE, an architecture that packs 512 experts into one-quarter the space. It achieves this by compressing tokens before routing them through expert layers.

The breakthrough targets memory bandwidth, not compute speed. Could optimizing bytes moved instead of FLOPs redefine production MoE economics?

In today's BrainScriblr:

  • NVIDIA's LatentMoE fits 512 experts in 25% space
  • TypeSafe's Jev returns typed decisions, not text
  • Google AX security gap leaves ports unenforced
  • Hugging Face hires MLX creator Jun Kim

NVIDIA's LatentMoE: Running 512 Experts in a Quarter of the Space

The Scoop: NVIDIA published LatentMoE, a mixture-of-experts architecture that fits 512 experts inside a quarter of the space.

It lifts MMLU-Pro scores 5.65 points with the same active parameters.

The Technical Details:

  • Projects tokens down 4× before expert dispatch: from model dimension 4,096 to latent dimension 1,024 before routing payload.
  • Performs all expert computation in latent space: 512 experts of width 1,024 instead of 128 at 4,096, then projects back up.
  • Router reads full-dimension tokens: routing decisions happen at d=4,096, only the routed payload compresses—preserving routing quality.
  • Activates 24 of 512 experts per token: same bytes per token (132 MB expert weights, 60 KB all-to-all) as baseline's 6-of-128.
  • Shipped in Nemotron 3 Super (March 2026, 120B total/12B active) and Ultra (June 2026, 550B total/55B active).

Why It Matters for You: Production MoE serving is memory-bound, not compute-bound. Interactive deployments spend 95%+ of latency streaming expert weights from HBM. Throughput workloads burn bandwidth on all-to-all token routing across GPUs. LatentMoE cuts both dominant costs by 75% without sacrificing model quality. The 95B model achieves baseline accuracy at 33 MB expert weights per token.

Compression enables 10³¹ times more expert combinations at iso-cost. Standard MoE scaled to match accuracy requires 350B additional parameters. NVIDIA projects 3.5× throughput gains at trillion-parameter scale when decode-heavy workloads dominate.

The Bigger Picture: MoE research optimized FLOPs because that's what academic benchmarks measure. Production bills optimize bytes moved because that's what GPUs wait on.

LatentMoE is the first major architecture to treat memory bandwidth as the primary constraint. It follows DeepSeek's MLA approach to KV cache compression—find what's expensive, shrink it, reinvest the savings.


Google's AX Agent Orchestrator Has a Security Gap: Ports Are Decorative

The Scoop: Google's AX agent orchestrator accepts port numbers in egress rules. None of them actually get enforced at runtime.

The Technical Details:

  • AX v0.3.0 defines `HostRule` with host and port fields in `pkg/apis/v1alpha1/ax.proto`.
  • The sandbox API wire format strips port data before reaching enforcement.
  • Six egress rules in the repo specify ports like `api.openai.com:443`.
  • The CLI reads your port declarations and prints them back.
  • No downstream component validates or enforces the port you declared.

Why It Matters for You: Production agent deployments pinning traffic to specific ports have no actual enforcement. Your security posture looks stricter in manifests than in runtime. Agents can reach any port on an allowed host. Budget extra engineering time to audit your actual network boundaries. Competitors using Temporal or Prefect may have tighter sandbox controls.

The Bigger Picture: This gap mirrors Kubernetes early days when NetworkPolicy objects existed before CNI plugins. Control-plane declarations need runtime teeth or they become documentation theater.


TypeSafe's Jev: A Model That Returns Decisions, Not Text

The Scoop: TypeSafe AI released Jev, a model that returns typed decisions, not text. It claims 100x speed and 200x cost improvements over LLMs for agent decisions.

The Technical Details:

  • Jev skips autoregressive token generation entirely with a single parallel pass.
  • Trained with RLCD (Reinforcement Learning for Calibrated Decisions), not RLHF.
  • Returns typed outputs with calibrated confidence scores for every decision.
  • Input pricing is $42 per billion tokens, 238x cheaper than Claude.
  • Eliminates JSON parsing and validation errors by construction, not post-processing.

Why It Matters for You: Most agent control flow decisions don't need full LLM reasoning. Jev targets binary choices: retry this API call? Which tool next? TypeSafe founder Diogo Almeida (ex-OpenAI RLHF researcher) raised $40M from DCVC. The model can't produce malformed output, cutting debugging time.

The Bigger Picture: This follows the pattern of specialized models replacing general-purpose ones. Just as BERT encoders beat GPT-3 for classification, decision models may dominate agent workflows.


Hugging Face Hires MLX Creator Jun Kim to Accelerate Local AI on Apple Silicon

The Scoop: Hugging Face hired Jun Kim, creator of oMLX, to support Apple's MLX framework full-time. His wrapper library graduates from side project to fully funded development.

The Technical Details:

  • oMLX provides a high-level wrapper around Apple's base MLX framework for easier integration.
  • The library stays Apache 2.0 licensed with Jun continuing as lead maintainer.
  • Plans include faster transitions from transformers definitions to reference MLX implementations across engines.
  • Collaboration targets mlx-lm and mlx-vlm teams (Cheng, Prince, Yagil) for upstream contributions.
  • Focus on building blocks for local AI that work across LMStudio and other tools.

Why It Matters for You: This hire signals maturation of Apple Silicon AI infrastructure beyond hobby projects. Organizations running models on Mac hardware gain stability and faster feature development. The streamlined transformers-to-MLX pipeline reduces deployment time for new models on Apple devices. Hugging Face positioning itself as the central hub strengthens its ecosystem lock-in.

The Bigger Picture: Local AI infrastructure mirrors the shift from mainframes to personal computing. Companies increasingly demand hybrid strategies that balance cloud costs with on-device capabilities.


📡 AI Discoveries

1. AI Models Achieving Autonomous Self-Improvement as Anthropic's Claude Leads 26% of R&D Anthropic's Claude AI is now autonomously handling over a quarter of the company's model research and development, marking a critical milestone toward AI systems that can improve themselves with minimal human supervision. This represents a major leap toward recursive self-improvement capabilities that could accelerate AI advancement exponentially. — ABC News, 2026-09-21

2. OpenAI Proposes Global AI Safety Standards for Recursive Self-Improvement OpenAI has published proposals for international standards governing frontier AI safety, specifically addressing recursive self-improvement capabilities, following warnings from researchers like Jacob Coxon who claimed companies are 'gambling with our lives.' This marks a significant industry response to growing concerns about AI existential risks. — CNBC, 2026-09-21

3. AI Co-Scientists Revolutionizing Research Methodology Across Scientific Fields AI systems are now being integrated as co-scientists in research teams, fundamentally transforming how scientific research is conducted. This development raises important questions about research attribution and credit in the AI era, particularly following OpenAI's recent mathematics breakthrough. — Nature, 2026-09-18


🌍 AI for Good

1. Gates Foundation Builds Coalition to Extend AI's Humanitarian Benefits to Underserved Communities The Gates Foundation is accelerating AI development for humanitarian applications, specifically focusing on extending benefits to poor communities currently 'shut out' from the technology, while balancing calls from major AI companies to slow advanced model development. — Greenwich Time, 2026-09-15

2. Major Initiative Launched to Ensure AI Speaks Everyone's Language for Global Equity A joint commitment from the Gates Foundation focuses on making AI accessible across all languages, ensuring that AI's problem-solving capabilities reach people most often left out and can help address the world's toughest challenges in underserved communities. — Gates Foundation, 2026-09-20

3. Research Reveals AI Adoption Failing to Close Accessibility Gap, Urges Ethical Guardrails New research from Level Access highlights that despite broad AI adoption, accessibility gaps persist, emphasizing that strong AI guardrails are essential to keeping systems safe, trustworthy, and aligned with human values as AI becomes embedded in daily life. — PR Newswire, 2026-09-18


Partner Spotlight

Support BrainScriblr while discovering powerful AI tools (affiliate links):

  • n8n — No-code automation platform for AI workflows
  • Hume AI — Emotional intelligence API for human-centered AI
  • Railway — Cloud platform for deploying AI applications
  • Cudo Compute — Distributed cloud computing for AI workloads

Worth Your Inbox

Discover more quality AI and tech content:

  • SemiVision — Semiconductor industry insights and AI chip developments
  • Turing Post — Deep technical analysis of AI research and breakthroughs
  • FinOps Weekly — Cloud cost optimization and financial operations
  • CoreUpdates — Essential tech updates and startup intelligence
  • The Multiverse School — Learning and development in the AI era
  • Simple AWS — Practical AWS tutorials and cloud architecture
  • EarthConscious — Sustainable living and environmental consciousness

Subscribe to Brainscriblr

Broader AI commentary. Subscribers get new posts by email a day before they go live on the site.

Email signup is coming soon — in the meantime, follow the Brainscriblr RSS feed.

Need a Custom MCP System?

Configuration & integration for your stack — from tool selection to production deployment. The directory recommends. The consultancy configures.

Get Started →