Which company has the best AI Agent end of August?
Open32as of
Market pricing makes Anthropic the favorite at 99% across 32 tracked outcomes, as of .
What are the odds right now?
Each outcome shows its current market price — the market-implied probability it happens. Click an outcome for its full market page.
What moved the odds?
18 AI-matched news signals available across this event, newest first. Entry price is the called side when the news hit; the signed return marks it against the outcome’s latest price as of .
Last week, @SpaceXAI released Grok 4.6, its new frontier model designed for coding, agentic tasks, and knowledge work. • • Independent evaluations suggest that SpaceXAI has returned to the AI frontier competition based on both performance and cost. • • The economics of Grok 4.6 seem particularly compelling for AI agents. Artificial Analysis estimates that Grok 4.6 costs $0.84 per task on its Intelligence Index, placing it on the Pareto frontier for intelligence-versus-cost. • • More from @MattarARK in this week's newsletter
DEEPSEEK CHALLENGES ANTHROPIC WITH NEW AI MODEL • • DeepSeek has unveiled an experimental multimodal AI model capable of analyzing images, screenshots and text while performing autonomous tasks. • • Called DeepSeek-V4-Flash-Vision-Exp, the model reportedly approaches Anthropic’s Opus 4.8 in multimodal agent performance while retaining DeepSeek’s latest text, reasoning and agent capabilities. • • The release highlights intensifying U.S.-China AI competition, with Chinese developers increasingly offering advanced models at lower costs.
Congrats to @deepseek_ai for their release of DeepSeekv4 Pro 0813 1.5T It massively beats Nemotron3 Ultra on agentic tasks. • • This is in addition to DeepSeekv4 Flash 0731 beating Nemotron3 Ultra massively too, while having 4.2x fewer active parameters and close to 2x fewer total parameters! • • There are lots of brilliant people working on Nemotron, but fundamentally, committee-based model frontier development does not work. 1/3
NVIDIA $NVDA JUST RELEASED NEMOTRON 3.5 LIGHTNING, ITS MOST EFFICIENT MODEL YET FOR LONG-RUNNING AI AGENTS • • The 30-billion-parameter mixture-of-experts model is built for specialized tasks inside larger multi-agent systems, following the earlier Nemotron 3 Nano release. • • Alongside it, NVIDIA released NeMo Switchyard, an open source library that intelligently routes each request to the most capable model for the job, without developers having to rewrite their applications.
NVIDIA Nemotron 3.5 Lightning Delivers Fast, Accurate Specialized Task Execution for Long-Running Agents. • Long-running AI agents spend most of their time on high-volume execution: tool calls, result validation, and subagent delegation. Using a frontier reasoning...
Congrats to @deepseek_ai on their release of DeepSeekv4 Flash 0731 It massively beats Nemotron3 Ultra on agentic tasks while having 4.2x fewer active parameters and close to 2x fewer total parameters! • • Committee-based model frontier development does not work. A focused team is what matters
A few thoughts on the $META AI model's progress, because I think it is significant. • • 1. It does seem that $META has now leapfrogged $GOOGL in model quality when it comes to Muse 1.2 for many use cases, which is very surprising given the timeframe. • • 2. This is still the "Muse Spark" family of models; $META already told us that bigger and more capable models (Watermelon) are coming. Given Muse Spark 1.2's performance already, the Watermelon family should be in the Fable category, which is impressive. • • 3. It does seem like $META is finally on a product scaling velocity curve (so the scaling foundations of their lab are set after 1 year of overhaul). And the shipping velocity is very good (3 releases in 4 months). • • 4. Given the recent rumored price hike of DeepSeek (if you use DeepSeek directly), it is clear that having enough compute to serve customers is critical. It doesn't help you if you have a great model, but most can't use it. $META is one of the few companies with compute capabilities comparable to, if not larger than, those of Anthropic or OpenAI.
Meta Unveils First AI Coding Agent 'Muse Code'... Challenging OpenAI and Antropic
Meta launches its first AI agent, aiming to challenge Anthropic and OpenAI.
Meta Unveils First AI Coding Agent… "Competing on Price, Not Performance"
Meta just shipped the beta version of its first coding agent. • • Muse Code is the agent system around Muse Spark 1.2, adding planning, tools, persistent session context, parallel sub-agents, and automatic validation. • • So it can handle one large software task for hours without needing a human to prompt every next step. • • - Scores 82.9% on Terminal-Bench 2.1 against 86.7% for Claude Code running Opus 5. • • - On DeepSWE 1.1 it lands at 59.3%, behind Opus 5 (65.0%) and GPT 5.6 Terra (64.8%). • • - installs with one command • • - You get background agents that stay alive for the whole session and accumulate context, instead of restarting cold on every task. • • - When a job grows big enough, the work fans out to sub-agents running in parallel inside isolated worktrees. • • - Every model call, tool run and edit hits a local event log before it executes, so a crash resumes from the last entry with nothing re-prompted. • • - That durability is why it could run 1,000+ tool calls over 24 hours on NVIDIA Hopper and keep finding kernel improvements deep into the session. • • - Pricing matches Muse Spark 1.1: $1.25/$4.25 per million input/output tokens
Meta Seeks To Take on OpenAI and Anthropic With a New AI Agent For Software Developers. • Meta is making a major push into AI-powered software development with the debut of Muse Code, its first dedicated coding agent.
*Meta Platforms Releasing Coding Agent to Compete With OpenAI, Anthropic, Company Says -- WSJ
$META | Meta Debuts First AI Coding Agent To Take On Anthropic And OpenAI - @CNBC •
Alphabet, Google's parent company, announced a major leadership overhaul of its artificial intelligence division. Demis Hassabis, head of AI, stepped down from his top management role, and several Gemini model leaders, including veteran engineer Jeff Dean, also left the company. This radical change comes at a pivotal moment for Google's DeepMind, as the main version of its latest Gemini model has yet to be released despite a planned June launch, raising concerns among investors and the industry that Google is falling behind rivals Anthropic and OpenAI. Alphabet shares fell 4% following the announcement of the changes. #Arabic_Business
Has China truly begun to surpass America in artificial intelligence? The question may seem shocking, but current events are prompting markets to ask it. Alibaba unveiled its Qwen 3.8 Max model, claiming it outperforms Anthropic's most powerful models in some tests. Months ago, many predicted that US chip restrictions would slow China down, but the opposite occurred: DeepSec, then Kemei, and now Alibaba, all at a surprisingly rapid pace. Alibaba's stock jumped more than 7% after the announcement. If China is advancing this quickly despite the restrictions, what would happen if they disappeared one day?
#Arabic_Business
Alibaba Unveils Qwen3.8-Max Model—China’s Latest AI Challenger To OpenAI And Anthropic. • The Chinese tech giant hailed its model as its “most capable.”
What is this event about?
This market will resolve according to the company that owns the model that has the highest rank based on the arena.ai Agent Arena Leaderboard when the table under "Agent Arena" filtered for "Models" is checked on August 31, 2026, 12:00 PM ET.
Results from the "Rank" column under the "Agent Arena" Leaderboard tab at https://arena.ai/leaderboard/agent filtered for "Models" will be used to resolve this market.
Models will be ordered primarily by their leaderboard rank at the market’s check time. If two or more models are tied on rank, they will be ordered by which model is listed higher on the leaderboard. If a tie still remains, alphabetical order of company names as listed in this market group will be used as a final tiebreaker (e.g., “Google” would be ranked ahead of “SpaceXAI”). This market will resolve based on the company that occupies first place under this ranking.
The resolution source for this market is the arena.ai Agent Arena Leaderboard. If this resolution source is unavailable at check time, this market will remain open until the leaderboard comes back online and will resolve based on the first check after it becomes available. If it becomes permanently unavailable, this market will resolve to "Other".