A
Artificial Analysis@ArtificialAnlys·22m ago
🚀 Release

Xiaomi open-weights MiMo-V2.6-Pro: AA Intelligence Index 46, tops open weights on cost frontier

Xiaomi open-weights MiMo-V2.6-Pro: AA Intelligence Index 46, tops open weights on cost frontier

@ArtificialAnlys reports Xiaomi released open-weights MiMo-V2.6-Pro: Artificial Analysis Intelligence Index 46 (up from 26 on V2.5-Pro), now the top open-weights model, landing on the intelligence-vs-cost Pareto frontier at about $0.13 per Index task with unchanged token prices ($0.435/M input with 99% cache-hit discount, $0.87/M output). It is a 1.02T-total / 42B-active MoE.

⚡ Key Takeaways
  • Open weights: AA Intelligence Index 46, big jump from V2.5-Pro (26) and tops open-weights.
  • Cost: ~$0.13/task; $0.435/M input (99% cache-hit discount) and $0.87/M output unchanged.
  • Architecture: 1.02T-total / 42B-active MoE; more AA analysis forthcoming.
Read details
C
Cognition@cognition·1h ago
🛠️ Tooling

Cognition ships Devin Cloud in Terminal and devin ssh: CLI cloud sessions plus SSH into Devin VMs

Cognition ships Devin Cloud in Terminal and devin ssh: CLI cloud sessions plus SSH into Devin VMs

@cognition introduced Devin Cloud in Terminal and devin ssh: create, steer, and resume Devin Cloud sessions from the CLI with /cloud, and for the first time SSH into Devin’s dedicated VM then /handoff the work back to your own device.

⚡ Key Takeaways
  • New CLI: /cloud to create, steer, and resume Devin Cloud sessions in the terminal.
  • devin ssh: first-time SSH into Devin’s dedicated VM.
  • /handoff: return cloud work to your own device for local workflows.
Read details
T
Teknium@Teknium·2h ago
🛠️ Tooling

Hermes Agent official plugin: Claude SDK restores Claude Code subscriptions inside Hermes

Hermes Agent official plugin: Claude SDK restores Claude Code subscriptions inside Hermes

@Teknium (Nous Research) shipped a new official Hermes Agent plugin that uses the Claude SDK without prior tradeoffs so Claude Code subscriptions work inside Hermes Agent again; install docs at hermes-agent.nousresearch.com/docs/plugins/claude-subscription-directsdk.

⚡ Key Takeaways
  • Official plugin: Claude SDK integration without prior tradeoffs.
  • Restores Claude Code subscription use inside Hermes Agent.
  • Install via Hermes docs: claude-subscription-directsdk plugin page.
Read details
A
Artificial Analysis@ArtificialAnlys·3h ago
📊 Benchmark

Artificial Analysis: Grok 4.7 + Grok Build hits Coding Agent Index 56 (4th); Intelligence Index 46 puts SpaceXAI in top 4 labs

Artificial Analysis: Grok 4.7 + Grok Build hits Coding Agent Index 56 (4th); Intelligence Index 46 puts SpaceXAI in top 4 labs

@ArtificialAnlys evaluated Grok 4.7 (xhigh): Intelligence Index 46 (+2 vs 4.6), putting SpaceXAI in the top 4 labs. With Grok Build it scores 56 on the Coding Agent Index (+9 vs 4.6), 4th among native harnesses behind Claude Fable 5.1, GPT-6 Astra, and Claude Opus 5. Pricing stays $2/$6 per 1M in/out ($0.50 cache hits), 500k context.

⚡ Key Takeaways
  • Coding Agent Index: Grok 4.7 + Grok Build scores 56 (+9), 4th among native harnesses.
  • Intelligence Index 46 (+2) puts SpaceXAI in the top 4 labs; +111 Elo on AA-Briefcase.
  • Same $2/$6 per 1M pricing as 4.6 ($0.50 cache hits), 500k context; gains come with higher token use.
Read details
ADSponsored
@
@elonmusk@elonmusk·4h ago
🔥 Trending

Grok 4.7 combines intelligence, speed, and low cost

Grok 4.7 combines intelligence, speed, and low cost

Elon Musk amplified SpaceXAI’s Grok 4.7 release. The model works longer on difficult tasks, checks its work more carefully, and ships with the strongest safeguards to date. He called it a strong mix of intelligence, speed, and low cost.

⚡ Key Takeaways
  • Elon Musk amplified SpaceXAI’s Grok 4.7 release. The model works longer on difficult tasks, checks its work more carefully, and ships with the strongest safeguards to date. He called it a strong mix of intelligence, speed, and low cost.
Read details
@
@elonmusk@elonmusk·4h ago
🚀 Release

Elon Musk: Grok 4.7 combines intelligence, speed, and low cost

Elon Musk: Grok 4.7 combines intelligence, speed, and low cost

Elon Musk quote-posted SpaceXAI, calling Grok 4.7 a strong combination of intelligence, speed, and low cost, and also posted “Try Grok 4.7.” The quoted SpaceXAI post says Grok 4.7 works longer on difficult tasks, checks its work more carefully, and ships with their strongest safeguards to date.

⚡ Key Takeaways
  • Elon Musk quote-posted SpaceXAI, calling Grok 4.7 a strong combination of intelligence, speed, and low cost, and also posted “Try Grok 4.7.” The quoted SpaceXAI post says Grok 4.7 works longer on difficult tasks, checks its work more carefully, and ships with their strongest safeguards to date.
Read details
@
@elonmusk@elonmusk·4h ago
🔥 Trending

Elon Musk: Grok 4.7 combines intelligence, speed, and low cost

Elon Musk: Grok 4.7 combines intelligence, speed, and low cost

Elon Musk quote-posted SpaceXAI, calling Grok 4.7 a strong combination of intelligence, speed, and low cost. Official notes say it works longer on difficult tasks, checks its work more carefully, and ships with their strongest safeguards; it is live in Cursor, Grok Build, and the API.

⚡ Key Takeaways
  • Elon Musk quote-posted SpaceXAI, calling Grok 4.7 a strong combination of intelligence, speed, and low cost. Official notes say it works longer on difficult tasks, checks its work more carefully, and ships with their strongest safeguards; it is live in Cursor, Grok Build, and the API.
Read details
S
SpaceXAI@SpaceXAI·4h ago
🚀 Release

SpaceXAI ships Grok 4.7: notable step up from 4.6 at the same price and speed, live in Cursor, Grok Build, and API

SpaceXAI ships Grok 4.7: notable step up from 4.6 at the same price and speed, live in Cursor, Grok Build, and API

@SpaceXAI launched Grok 4.7 as a notable improvement over Grok 4.6 at the same price and speed. Official notes say it works longer on hard tasks, checks its work more carefully, and ships with their strongest safeguards; it is live in Cursor, Grok Build, and the API.

⚡ Key Takeaways
  • Official release: clear upgrade over Grok 4.6 at unchanged price and speed.
  • Runs longer on hard tasks, checks work more carefully, strongest safeguards yet.
  • Available now in Cursor, Grok Build, and the API for coding-agent workflows.
Read details
ADSponsored
l
lauren@poteto·5h ago
🛠️ Tooling

Ex-Cursor engineer @poteto publishes Cursor Compile talk: shipping 2,500 production PRs in a month with coding agents

SpaceXAI engineer and ex-Cursor @poteto published the talk planned for Cursor Compile London: how she shipped about 2,500 production PRs in a month with coding agents. She missed the event while livestreaming Grok Bot Galaxy and released the talk for free on X.

⚡ Key Takeaways
  • Cursor Compile London talk released for free on X after a livestream conflict.
  • Core case: about 2,500 production PRs shipped in one month with coding agents.
  • Author is at SpaceXAI (ex-Cursor); talk targets Cursor and coding-agent workflows.
Read details
K
Kimi Developers@KimiDevs·8h ago
🚀 Release

Moonshot ships Kimi Code Desktop: multi-agent Mac/Windows app for parallel long-horizon coding

Moonshot ships Kimi Code Desktop: multi-agent Mac/Windows app for parallel long-horizon coding

@KimiDevs launched Kimi Code Desktop for Mac and Windows: a focused workspace to manage multiple agents, run tasks in parallel, and stay in sync on long-horizon programming work, aiming for faster and more reliable coding results.

⚡ Key Takeaways
  • Official desktop client now available on macOS and Windows.
  • Manage multiple agents in one workspace with parallel task runs.
  • Built for long-horizon coding with ongoing agent sync and faster task completion.
Read details
A
Avi Chawla@_avichawla·11h ago
🛠️ Tooling

Beacon open-sources cross-harness self-improving memory: Jev scores runs into reusable skills

@_avichawla covered Beacon by Asymptote Labs (github.com/Asymptote-Labs/agent-beacon): a cross-harness memory layer for Claude Code, Codex, Cursor, OpenCode, and 20+ more. It captures full session history, uses Jev to score evidence/reuse/human-correction signal, and promotes high-signal workflows into reusable skills so lessons transfer across harnesses.

⚡ Key Takeaways
  • Open-source github.com/Asymptote-Labs/agent-beacon: memory layer across 20+ coding harnesses.
  • Jev scores sessions on evidence, reuse potential, and human corrections (promote/review/discard).
  • High-signal workflows become reusable skills transferable across Claude Code, Codex, Cursor, OpenCode.
Read details
A
Akshay@akshay_pachaar·13h ago
🛠️ Tooling

HarnessRouter open-sources Unified Harness Protocol: one API for Codex, Claude Code, Hermes, and more

@akshay_pachaar highlighted HarnessRouter (github.com/HarnessRouter/harnessrouter, Apache-2.0): an “OpenRouter for agent harnesses” implementing the Unified Harness Protocol (UHP) so Codex, Claude Code, Hermes, DeepSeek Harness, System One (Jev), and 9+ other harnesses share one API for sessions, streaming, files, cancellation, and failure handling—self-hosted with your own keys.

⚡ Key Takeaways
  • Open-source github.com/HarnessRouter/harnessrouter (Apache-2.0); site harnessrouter.ai.
  • Unified Harness Protocol: one API for sessions, streaming, files, cancellation, and failures.
  • Supports Codex, Claude Code, Hermes, DeepSeek Harness, System One (Jev), and 9+ other harnesses.
Read details
l
lifcc@mylifcc·14h ago
🚀 Release

China Telecom open-sources Xing4.0-29B-A4B: 29B/4B MoE, SWE-bench Verified 75.0, Ascend-trained

China Telecom open-sources Xing4.0-29B-A4B: 29B/4B MoE, SWE-bench Verified 75.0, Ascend-trained

@mylifcc highlighted China Telecom AI’s open-source Xing4.0-29B-A4B (TeleChat lineage; HF XingChen-AGI/Xing4.0-29B-A4B, Apache-2.0): 29B total / 4B active per token with native 256K context. The model card reports SWE-bench Verified 75.0 and Terminal-Bench 2.1 57.50, claims agent-framework alignment (OpenCode / Claude Code), and was trained end-to-end on Ascend NPUs with MindSpore.

⚡ Key Takeaways
  • HF XingChen-AGI/Xing4.0-29B-A4B (Apache-2.0): 29B/4B MoE, 256K context.
  • Card scores: SWE-bench Verified 75.0; Terminal-Bench 2.1 57.50.
  • Trained on Ascend+MindSpore; claims OpenCode / Claude Code agent-framework alignment.
Read details
D
Diogo Almeida@CompleteSkeptic·15h ago
🛠️ Tooling

TypeSafe co-founder shares coding-agent design notes: routing, compaction, subagents, skills, and MCP

@CompleteSkeptic (Diogo Almeida, TypeSafe; RLHF background) published public design notes on TypeSafe × coding agents (Google Doc), covering routing, context compaction, subagents, skills, and MCP tradeoffs, and invited the community to experiment. The post drew ~930 likes / ~56k views the same day.

⚡ Key Takeaways
  • Public Google Doc on coding-agent routing, compaction, subagents, skills, and MCP.
  • Authored by a TypeSafe principal; explicitly invites community experimentation.
  • Same-day high engagement (~930 likes / ~56k views) as a design reference.
Read details
J
James Grugett@jahooma·16h ago
📉 Price Cut

Freebuff launches $8/mo Starter: ad-supported high-usage coding plan centered on DeepSeek V4.1 Flash / GLM Flash

Freebuff launches $8/mo Starter: ad-supported high-usage coding plan centered on DeepSeek V4.1 Flash / GLM Flash

@jahooma (Freebuff/Codebuff) announced an $8/mo Starter tier: the post claims ~15 hours/day of DeepSeek V4.1 Flash with unlimited messages because ads subsidize usage. freebuff.com/plans currently lists Starter at $8/mo with about 10 hrs/day DeepSeek V4.1 Flash or 30 hrs/day GLM 5.3 Flash across Web, Desktop, and CLI.

⚡ Key Takeaways
  • Founder announced $8/mo Starter: ad-supported, claiming ~15h/day DeepSeek V4.1 Flash with unlimited messages.
  • Plans page lists Starter at $8/mo with ~10h/day DeepSeek V4.1 Flash or ~30h/day GLM 5.3 Flash.
  • Works across Web, Desktop, and CLI as a low-cost alternative tier to Claude Code / Cursor / Codex.
Read details
v
vechen@miu21590·18h ago
🛠️ Tooling

Mid-task Jev in Codex raises/lowers GPT-6 Astra reasoning effort; author reports ~50% Astra cost cut without breaking prompt cache

Mid-task Jev in Codex raises/lowers GPT-6 Astra reasoning effort; author reports ~50% Astra cost cut without breaking prompt cache

@miu21590 demoed using Jev inside a running Codex session to raise/lower GPT-6 Astra reasoning effort—more thinking when stuck, less on routine steps—reporting ~50% lower Astra costs and faster runs without breaking prompt caching. Versus the usual pre-task Jev model pick, this wires the decision model mid-run for cost control.

⚡ Key Takeaways
  • Call Jev mid-Codex-run to raise/lower GPT-6 Astra reasoning effort.
  • Author tests: ~50% lower Astra cost, faster, prompt cache intact.
  • Moves Jev from pre-task model pick to in-task cost control.
Read details
Z
ZCode@zcode_ai·19h ago
🚀 Release

Zhipu open-sources ZCode and ships v3.14.0 after repo-snapshot security fix; CAICT/NSFOCUS confirm OSS emptied

Zhipu open-sources ZCode and ships v3.14.0 after repo-snapshot security fix; CAICT/NSFOCUS confirm OSS emptied

@zcode_ai addressed community-reported product security issues by open-sourcing ZCode (github.com/zai-org/ZCode, Apache-2.0 — desktop/Web/Agent CLI) and shipping client v3.14.0 that removes Repo Wiki and the local-repo snapshot upload path. CAICT and NSFOCUS assessments confirm the zcode-prod Alibaba Cloud OSS bucket is zero-data and deleted; the team says referenced code data was not retained or used for training and will run an ongoing vuln-reporting reward process.

⚡ Key Takeaways
  • Open-sourced github.com/zai-org/ZCode (Apache-2.0): desktop, Web, and Agent CLI/runtime.
  • v3.14.0 removes Repo Wiki and local-repo snapshot generate/upload paths.
  • CAICT/NSFOCUS confirm zcode-prod OSS zero-data and deleted; no retention/training use claimed; vuln rewards coming.
Read details
J
Jaana Dogan ヤナ ドガン@rakyll·19h ago
🛠️ Tooling

Agent Substrate opens env: no-code environment abstraction with built-in MCP fs and shell tools

Agent Substrate opens env: no-code environment abstraction with built-in MCP fs and shell tools

@rakyll (Jaana Dogan) announced github.com/agent-substrate/env (Apache-2.0): an isolated, snapshottable environment layer on Agent Substrate with ate-env CLI/API and guest daemon, plus a built-in MCP server exposing filesystem and shell tools so any MCP client can try Substrate without writing integration code first (API still alpha).

⚡ Key Takeaways
  • Open-sourced agent-substrate/env: CLI, API, guest daemon, Go/Python clients.
  • Built-in MCP for filesystem and shell — try with any MCP client.
  • Substrate container actors: snapshot, schedule, multiplex idle envs; API still alpha.
Read details
A
Average AI Bro@AverageAiBro·20h ago
🛠️ Tooling

retrieval-mcp 0.1.6: Rust MCP whose four-tool default cuts ~24–42% code-search input tokens

@AverageAiBro highlighted retrieval-mcp (github.com/abendrothj/retrieval-mcp) v0.1.6, a Rust stdio MCP for coding agents whose measured four-tool default (search_exact, read_source, find_callers, search_concept) cuts schema tax; README claims ~24–42% fewer input tokens vs specialist search MCP with matched answer quality. Install via cargo or the release binary, then register with Claude Code/Codex.

⚡ Key Takeaways
  • Default four tools: search_exact, read_source, find_callers, search_concept.
  • README: ~24–42% fewer input tokens vs specialist search MCP, matched quality.
  • Rust stdio MCP; cargo install or v0.1.6 binary for Claude Code/Codex.
Read details
D
Dan Kornas@DanKornas·20h ago
🛠️ Tooling

Interceptor: CLI and MCP that drive your signed-in Chrome/Brave/Safari session and macOS apps

Interceptor: CLI and MCP that drive your signed-in Chrome/Brave/Safari session and macOS apps

@DanKornas highlighted Interceptor (github.com/Hacker-Valley-Media/Interceptor), an agent browser and macOS control stack: one CLI plus extension drives your existing Chrome/Brave/Safari session (cookies, logins, tabs intact), with page actions, passive network capture, record-and-replay, and macOS accessibility/trusted input plus on-device vision/speech. Wires into Claude Code, Codex, Gemini CLI, and Cursor.

⚡ Key Takeaways
  • Drive existing Chrome/Brave/Safari sessions instead of a blank automated profile.
  • Passive capture of fetch/XHR/SSE/WebSocket traffic plus record-and-replay.
  • macOS app control; wires Claude Code, Codex, Gemini CLI, and Cursor.
Read details