ChatGPT Work's cloud computer environment can now sign in to websites on web and mobile for Plus, Pro, and Business users, storing credentials inside isolated browser forms without exposing passwords to the model.
Key Takeaways
Authenticated computer-use unlocks high-value workflows behind logins that were previously unreachable;
Secure browser credential isolation ensures the underlying LLM never sees or ingests plaintext passwords;
Simultaneous web and mobile rollout positions ChatGPT Work as an autonomous workplace operator.
OpenAI released InferenceX benchmarks for JalapeΓ±o, its first custom inference silicon, showing 1.5β1.9x higher work per watt and 1.7β3.6x lower end-to-end latency across GPT-OSS 120B, DeepSeek R1, and Kimi K2.5.
Key Takeaways
OpenAI officially expands into first-party silicon, publishing public model-agnostic benchmarks across open weights;
Designed around agentic serving, unified architecture handles both prefill and decode stages simultaneously;
AI-assisted chip development achieved tapeout in 9 months and enabled rapid Codex-driven model porting in 2 months.
Google DeepMind released Gemini 3.7 Flash, delivering massive gains in agentic coding and debugging (FrontierCode 43.6%, DeepSWE 65.3%) with Day-0 availability in Antigravity and AI Studio at introductory rates.
Key Takeaways
Substantial accuracy gains over 3.6 Flash on complex debugging, issue resolution, and single-shot UI layouts;
Benchmark scores: FrontierCode 1.1 jumps to 43.6% (vs 34.4%) and DeepSWE v1.1 to 65.3% (vs 49.0%);
Day-0 integration in Antigravity, AI Studio, and Android Studio at promotional $0.75 / $3.75 per 1M rates.
Alibaba Qwen released Qwen3.8-Flash, an open-weights 125B MoE model (6B active) previewing the Qwen4 architecture with 1/9 training cost, 262K native context, and ultra-cheap API pricing.
Key Takeaways
New core architecture combining GDN gated residuals, sparse attention, N-gram embeddings, and Muon optimizer;
Coding & SWE benchmarks: 58.7 on DeepSWE 1.1, 62.5 on SWE-bench Pro, and 73.9 on CoWorkBench;
Day-0 support across QwenCloud, OpenRouter, vLLM, SGLang, and Unsloth (runnable locally with ~75GB RAM via GGUF).
DeepSeek open-sources its Multi-Token Prediction (MTP) agent acceleration framework. By speculatively predicting syntax blocks in parallel, multi-file code diff and test generation in Cline and Aider achieves over 3x speedup with zero degradation in AST validity.
Key Takeaways
DeepSeek releases Multi-Token Prediction (MTP) framework tailored for coding agents;
Delivers 300% throughput acceleration during multi-file refactor and test suites;
Directly compatible with local vLLM and Ollama deployments.
OpenRouter confirmed the viral mystery Ox Alpha model is Zhipu's GLM-5.3-Flash, processing over 20 trillion tokens in six days as the fastest-growing model in router history, now live with 1M context.
Key Takeaways
Processed 20T tokens in six days, setting an all-time adoption record for new models on OpenRouter;
Hybrid sparse-linear attention delivers ultra-cheap long-horizon coding and agent inference across 1M context;
Promotional launch rates set at $0.075 / $0.25 / $0.015 per million tokens (input / output / cache) through Sept 9.
xAI announced Grok 4.6, delivering major speed and reasoning enhancements tailored for long-running coding agents, with Day-0 availability in Grok Build, Cursor, and the API at $2 / $6 per million tokens.
Key Takeaways
Deeply optimized via RL for multi-file refactoring, dependency tracking, and autonomous agent loops;
Priced at $2/M input and $6/M output, keeping Grok 4.5 parity while boosting reasoning and speed;
Available Day-0 in Grok Build, Cursor, and the API with 2x usage allowance during launch week.
Perplexity introduced Portable Computer on NVIDIA DGX Spark, providing a fully local-first agent stack powered by on-device PPLX 27B where private documents never leave the machine by default.
DeepSeek launched deepseek-v4-flash-vision-exp on its API, maintaining ultra-low V4-Flash pricing while achieving multimodal coding and agent performance approaching Claude Opus 4.8.
Key Takeaways
Preserves ultra-cheap V4-Flash token rates for both text and vision multimodal inputs;
Achieves massive leaps on multimodal agent benchmarks, rivaling Claude Opus 4.8 on GUI and diagram reasoning;
Day-0 support added in DeepSeek Harness 0.1.1 for seamless agent orchestration and benchmarking.