Daily AI News | OpenAI publishes first benchmark results for its Jalapeño inference chip
# Daily AI News | OpenAI publishes first benchmark results for its Jalapeño inference chip
## Lead Story
### OpenAI publishes first benchmark results for its Jalapeño inference chip
OpenAI published the first results for its custom Jalapeño inference chip. Across workloads including GPT‑OSS 120B, DeepSeek R1, and Kimi K2.5, it reports 1.5–1.9 times more work per watt and 1.7–3.6 times lower latency, with production deployment planned by the end of 2026.
**AI take:** The announcement moves OpenAI further from buying accelerators toward owning its inference stack. These are vendor measurements, so production scale, yield, and real serving cost remain decisive.
## More News
### Claude Chat and Cowork begin sharing memory
Anthropic has connected memory across Claude Chat and Cowork for Free, Pro, and Max users. People can view, edit, or delete stored memories, with more cautious defaults for sensitive material.
**AI take:** This reduces repeated setup, while making review of stored memory more important.
### Keenable raises $26 million to build a web index for AI agents
Keenable raised a $26 million seed round led by Accel. The company says its web index for AI agents covers more than 100 billion documents and its API is already in production.
**AI take:** Index quality, freshness, and query cost will decide whether developers adopt it.
### NVIDIA introduces the Jetson Orin Nano 2 edge-AI computer
NVIDIA introduced Jetson Orin Nano 2 with 78 TOPS of AI performance, 8GB of memory, and an eight-core Arm CPU. It says inference performance doubles versus the predecessor, with availability planned for the first half of 2027.
**AI take:** The device raises the local-inference ceiling for entry robotics and edge systems. Adoption will still depend on price, supply, and the developer-kit ecosystem.
### India’s AM Intelligence orders 9,000 NVIDIA Vera Rubin systems
Bloomberg reported that Indian AI data-center company AM Intelligence ordered 9,000 NVIDIA Vera Rubin systems. The purchase gives the next-generation platform a major infrastructure customer before broad deployment.
**AI take:** The order accelerates India’s domestic AI-capacity buildout and adds pressure to Rubin supply planning. Delivery timing and activated capacity still need confirmation.
### Quantization-Aware Healing compresses GPT‑OSS 120B to about 60 billion parameters
Multiverse Computing published a Quantization-Aware Healing result that compresses GPT‑OSS 120B into an approximately 60-billion-parameter MXFP4 model. It says the derivative beats the BF16 source on seven of nine benchmarks while using roughly four times less weight memory.
**AI take:** If independently reproduced, healing after compression could lower large-model deployment barriers. The performance claim currently rests mainly on the releasing team’s evaluation.
### XCENA details its MX1 CXL computational-memory architecture at Hot Chips
XCENA detailed MX1 with up to 2TB of DDR5, SSD backing, and more than 1,000 RISC‑V cores for near-memory computing. It claims up to 4.7 times higher throughput on selected kernels and targets mass production by the end of 2026.
**AI take:** Server compatibility and software support will matter more than selected-kernel claims.
### Alibaba’s Meoo turns prompts into native Android and iOS apps
Alibaba Cloud’s Meoo demonstrated a workflow that generates native Android and iOS apps from natural-language requirements, including preview, device testing, packaging, and cloud-backend setup. It targets creators without a full mobile-development workflow.
**AI take:** Moving from web prototypes to installable native apps is a meaningful step for AI coding tools. Store review, maintenance, and security remain part of the real delivery cost.
### OpenAI launches an Admin plugin for ChatGPT Work and Codex
OpenAI launched an Admin plugin that lets workspace administrators inspect usage, members, groups, access, limits, and spending, and perform authorized management actions. It inherits existing admin permissions and can automate routine tasks.
**AI take:** The enterprise-AI control plane is becoming conversational too. Convenience should be paired with least privilege, approvals, and operation auditing.
### ByteDance formally launches Doubao Work for task execution
ByteDance formally launched Doubao Work with browser and computer control, continued cloud execution, and Feishu collaboration. The product is positioned to complete multi-step tasks rather than only generate content.
**AI take:** Reliability, permissions, and recovery from long-task failures will determine daily use.
### OpenRouter launches a unified video-generation API
OpenRouter released an asynchronous video API at `/api/v1/videos` for models including Seedance, Veo, and Wan, with job polling and result downloads. Its documentation says the endpoint does not currently support zero-data retention.
**AI take:** Retention limits, queue latency, failure rate, and price remain key comparisons.
### Perplexity and NVIDIA introduce the local-first Portable Computer agent
Perplexity and NVIDIA introduced Portable Computer on DGX Spark, using Qwen3.8‑27B locally for research, organization, and file tasks without cloud-token charges. With permission, it can escalate to cloud models, and support for RTX and Windows devices is planned.
**AI take:** Local-first execution with permissioned cloud escalation is a practical privacy, cost, and capability tradeoff. Hardware price, task success, and wider device support will determine adoption.
### A model-poisoning path is disclosed in NVIDIA NemoClaw deployments
SiliconANGLE reported that CVE‑2026‑65105 can combine a malicious webpage with DNS rebinding to reach an exposed, unauthenticated Ollama endpoint and alter model templates behind a NemoClaw agent. The risk centers on misconfigured local services.
**AI take:** Local models are not automatically safe once agents connect browsers, models, and tools. Restricted binding, authentication, and isolated tool permissions are the immediate defenses.
### The SEC probes AI hedge fund Situational Awareness
TechCrunch reported that the U.S. Securities and Exchange Commission is probing AI hedge fund Situational Awareness, which had previously come close to collapse after trading losses. The inquiry is a new regulatory stage after the operating crisis.
**AI take:** An AI theme does not remove conventional market risk; concentration and leverage still need transparent controls. The probe may influence compliance expectations for new AI-focused funds.
## Rumors