Which company has the best AI model on LiveBench (Coding) end of September?
Open32as of
Market pricing makes Anthropic the favorite at 58% across 32 tracked outcomes, as of .
What are the odds right now?
Each outcome shows its current market price — the market-implied probability it happens. Click an outcome for its full market page.
What moved the odds?
12 AI-matched news signals available across this event, newest first. Entry price is the called side when the news hit; the signed return marks it against the outcome’s latest price as of .
Google Launches New AI Model in 3 Weeks… “Improves Coding Accuracy and Lowers Prices”
Meta Unveils First AI Coding Agent 'Muse Code'... Challenging OpenAI and Antropic
Meta Unveils First AI Coding Agent… "Competing on Price, Not Performance"
Meta just shipped the beta version of its first coding agent. • • Muse Code is the agent system around Muse Spark 1.2, adding planning, tools, persistent session context, parallel sub-agents, and automatic validation. • • So it can handle one large software task for hours without needing a human to prompt every next step. • • - Scores 82.9% on Terminal-Bench 2.1 against 86.7% for Claude Code running Opus 5. • • - On DeepSWE 1.1 it lands at 59.3%, behind Opus 5 (65.0%) and GPT 5.6 Terra (64.8%). • • - installs with one command • • - You get background agents that stay alive for the whole session and accumulate context, instead of restarting cold on every task. • • - When a job grows big enough, the work fans out to sub-agents running in parallel inside isolated worktrees. • • - Every model call, tool run and edit hits a local event log before it executes, so a crash resumes from the last entry with nothing re-prompted. • • - That durability is why it could run 1,000+ tool calls over 24 hours on NVIDIA Hopper and keep finding kernel improvements deep into the session. • • - Pricing matches Muse Spark 1.1: $1.25/$4.25 per million input/output tokens
Meta to take on Anthropic's Claude and OpenAI's Codex with new coding agent
Meta Muse Spark 1.2 & Muse Code Beta • • - Muse Spark 1.2 is Meta's new coding model, reaching near frontier-level performance on benchmarks like TerminalBench 2.1 and DeepSWE 1.1. • - Muse Code is now in beta: a terminal coding agent that uses persistent and parallel AI agents to plan, code, test, and validate large software engineering tasks.
i never thought i’d see a meta ai model reach parity with fable and gpt 5.6 at coding but here we are • • muse spark 1.2 beats gemini 3.6 flash and grok 4.5 in deepSWE. this is metas second model release in <4 weeks. • • they truly pulled off a miracle and the single common denominator is compute • • zuck aggressively invested in data centers and its paid off. well done.
*Meta Platforms Releasing Coding Agent to Compete With OpenAI, Anthropic, Company Says -- WSJ
$META | Meta Debuts First AI Coding Agent To Take On Anthropic And OpenAI - @CNBC •
What is this event about?
This market will resolve according to the company which owns the model with the highest Coding score on LiveBench.ai when the leaderboard is checked on September 30, 2026, at 12:00 PM ET.
Results from the “Coding” column of the leaderboard at https://livebench.ai/#/?cats=Coding, with the latest available LiveBench release selected and the category set to “Coding,” will be used to resolve this market.
Models will be ranked according to the specified score, with higher scores ranked ahead of lower scores. If two or more models have exactly the same score as displayed on the leaderboard, the model with the lower listed "cost per successful task" will be ranked ahead. If a tie still remains, alphabetical order of company names as listed in this market group will be used as a final tiebreaker (e.g., if the two models are tied by exact score and cost per successful task, “Google” would be ranked ahead of “SpaceXAI”). This market will resolve based on the company that occupies first place under this ranking.
The resolution source for this market is the LiveBench leaderboard. If this resolution source is unavailable at check time, this market will remain open until the leaderboard comes back online and will resolve based on the first check after it becomes available. If it becomes permanently unavailable, this market will resolve to “Other.”