Entries may carry one or more status tags so readers can judge maturity at a glance:
๐ค Awesome AI Agents 2026
Snapshot 2026-08-03 23:56:39 UTC ยท version 1
Research document
๐ค Awesome AI Agents 2026
The definitive curated list of AI models, agent frameworks, tools, protocols, and resources for 2026 โ the year agents went mainstream and AI became infrastructure.
Covering foundation models, multimodal AI, agent protocols (MCP/A2A), coding agents, computer use, generative AI, and more.
๐ท๏ธ Status Legend
Entries may carry one or more status tags so readers can judge maturity at a glance:
- ๐ New โ Added in the last 60 days, still settling.
- ๐ฆ Archived โ Repository archived by its owner; preserved for historical reference, no further updates expected.
- ๐ค Stale โ No commits in 6+ months; project may still work but is no longer actively maintained.
- โ ๏ธ Unverified โ Recent submission with limited independent traction (low stars / no third-party adoption / sole-maintainer / submitted to many awesome lists in parallel). Listed for completeness, not endorsed โ vet before using.
- ๐จ๐ณ Chinese ecosystem โ Project from a mainland-China team or primarily targeting the Chinese market.
- ๐ฅ Hot โ GitHub stars grew >20% in the last 30 days; community momentum.
- โก Updated โ Received a notable release or major feature in the last 14 days.
- ๐งช Experimental โ Promising but not production-ready; use for R&D only.
- ๐ฐ Freemium โ Core functionality free; paid tiers for scale/advanced features.
- ๐ Audited โ Has undergone independent security audit or formal verification.
- ๐จ๐ณ China-first โ Optimized for Chinese language, regulation, or infra stack.
Foundation Models ยท Multimodal AI ยท Protocols ยท Frameworks ยท IDEs & Builders ยท Memory ยท Tools ยท Sandboxing ยท Security ยท RAG ยท Coding ยท Physical AI ยท Simulation ยท Benchmarks ยท Computer Use ยท Browser & Web ยท Voice ยท Personal ยท Mobile ยท Enterprise ยท Evaluation ยท Research Tools ยท Learning ยท Chinese Ecosystem ยท Compare ยท Notable 2026 ยท Timeline
๐ Start Here
New to AI agents? Follow this path:
- ๐ Understand โ what an agent actually is vs. a chatbot
- ๐บ๏ธ Find your scenario โ Scenario Guide
- ๐งฉ Copy a proven setup โ Stack Recipes
- ๐ Pick the right tool โ Compare Tables
- โ ๏ธ Avoid common mistakes โ Anti-Picks
Already building? Jump to:
- ๐ Latest additions (July 2026) โข ๐ก๏ธ Security โข ๐ฐ Cost comparison
Quick Navigation
| Category | Description | Count |
|---|---|---|
| ๐ง Foundation Models | Latest LLMs from OpenAI, Anthropic, Google, Meta, and 22+ providers | 195+ |
| ๐จ Multimodal & Generative AI | Image, video, audio, and music generation | 40+ |
| ๐ Agent Protocols | MCP, A2A, and interoperability standards | 15+ |
| ๐๏ธ Agent Frameworks | Libraries for building autonomous AI agents | 23+ |
| ๐ ๏ธ Agent IDEs & Visual Builders | Visual / low-code environments for designing agent flows | 8+ |
| ๐ง Agent Memory | Persistent memory and context management | 20+ |
| ๐ Tool & API Integration | Connecting agents to external services | 20+ |
| ๐ฑ Agent Economy & Marketplaces | Where agents pay, get paid, and discover services | 5+ |
| ๐งช Sandboxing & Compute Isolation | Secure runtimes for agent-generated code | 9+ |
| ๐ก๏ธ Agent Security | Prompt injection defense and guardrails | 16+ |
| ๐ RAG & Knowledge | Retrieval-augmented generation systems | 20+ |
| ๐ป Coding Agents | AI-powered software engineering | 45+ |
| ๐ค Physical AI | Humanoid robots, embodied AI, industrial automation | 35+ |
| ๐ฎ Simulation & World Models | Sim environments for training and stress-testing agents | 10+ |
| ๐ Benchmarks | Leaderboards tracking frontier capability | 20+ |
| ๐ฅ๏ธ Computer Use | Desktop automation and OS-level control | 10+ |
| ๐ Browser & Web Agents | Agents that drive real browsers | 15+ |
| ๐ฃ๏ธ Voice & Multimodal Agents | Voice-enabled conversational AI | 10+ |
| ๐ฑ Personal AI Agents | Productivity and daily life assistants | 20+ |
| ๐ฑ Mobile Agents | Phone-control agents (Android / iOS) | 10+ |
| ๐ข Enterprise Platforms | Enterprise-grade agent deployment | 30+ |
| ๐ Evaluation & Observability | Testing, monitoring, and benchmarking | 30+ |
| ๐ฌ AI Research Tools | Tools for AI/ML research and experimentation | 15+ |
| ๐ Learning Resources | Papers, courses, and tutorials | 20+ |
| ๐จ๐ณ Chinese AI Ecosystem | Major projects from China-based teams | 25+ |
| ๐ Compare | Side-by-side comparison tables | โ |
| ๐บ๏ธ Scenario Guide | 56 curated scenario-to-tool mappings | 56 |
| ๐ Stack Recipes | Curated multi-tool combinations | 8 |
| โ ๏ธ Anti-Picks | What NOT to use and why | 15 |
Contents
- ๐ง Foundation Models 2026
- ๐จ Multimodal & Generative AI
- ๐ Agent Protocols & Standards
- ๐๏ธ Agent Frameworks
- ๐ ๏ธ Agent IDEs & Visual Builders
- ๐ง Agent Memory
- ๐ Tool & API Integration
- ๐ฑ Agent Economy & Marketplaces
- ๐งช Agent Sandboxing & Compute Isolation
- ๐ก๏ธ Agent Security
- ๐ RAG & Knowledge
- ๐ป Coding Agents
- ๐ค Physical AI & Embodied Agents
- ๐ฎ Agent Simulation & World Models
- ๐ Benchmarks & Leaderboards
- ๐ฅ๏ธ Computer Use & Desktop Agents
- ๐ Browser & Web Agents
- ๐ฃ๏ธ Voice & Multimodal Agents
- ๐ฑ Personal AI Agents
- ๐ฑ Mobile Agents
- ๐ข Enterprise Agent Platforms
- ๐ Agent Evaluation & Observability
- ๐ฌ AI Research Tools
- ๐ Learning Resources
- ๐จ๐ณ Chinese AI Ecosystem
- ๐ Compare โ Side-by-Side Tables
- ๐บ๏ธ Scenario Guide โ What Should I Use Forโฆ
- ๐ Stack Recipes โ Curated Tool Combinations
- โ ๏ธ Anti-Picks โ What NOT to Use Forโฆ
- ๐ Notable Agent Projects of 2026
- ๐ 2026 AI Timeline
๐ง Foundation Models 2026
The latest large language models powering the AI ecosystem, organized by company. 60+ models from 20+ providers.
OpenAI
GPT-Live-1 / GPT-Live-1 mini - ๐ July 8, 2026. OpenAI's full-duplex conversational voice model replacing Advanced Voice Mode. ChatGPT-only โ not exposed as an API model; for programmatic realtime voice use
gpt-realtime-2.1, and for streaming transcriptiongpt-live-transcribe($0.017/min). Listens and speaks simultaneously (zero turn-taking lag), handles interruptions, delegates complex queries to GPT-5.5 in the background while keeping the conversation flowing. GPT-Live-1 is default for paid users (Go/Plus/Pro); GPT-Live-1 mini is default for free users. Includes real-time live translation. Available on iOS, Android, and web.OpenAI Astra - ๐ โ ๏ธ Announced August 1, 2026 (public release date TBD). OpenAIโs next-generation model family announced in pre-release. An internal version reportedly resolved 10 unsolved mathematical and theoretical computer science problems in a single session (group theory, quantum complexity). 249-page manuscript + Lean 4 machine-verifiable certificates published publicly. โ ๏ธ No API access or public weights; announcement only; public release timeline unknown.
GPT-5.6 Sol - ๐ July 9, 2026 (GA; limited preview from June 26). OpenAI's frontier flagship in the GPT-5.6 family โ "Sol" is the most capable tier with advanced reasoning, coding, biology, and cybersecurity capabilities plus "max" reasoning and "ultra" sub-agent mode. Available on ChatGPT, Codex, and the OpenAI API. Launch was delayed briefly at US government request for a national-security review; rolled out to all users in stages following the trusted-partner preview.
GPT-5.6 Terra - ๐ July 9, 2026. Mid-tier model in the GPT-5.6 family offering GPT-5.5-parity performance at approximately 2ร lower cost. Designed for cost-efficient production workloads.
GPT-5.6 Luna - ๐ July 9, 2026. The fastest and most cost-efficient tier of GPT-5.6 โ optimised for high-volume, speed-critical tasks.
ChatGPT Work - ๐ July 9, 2026. OpenAI's agent that turns a goal into finished work โ acts across connected apps and files, stays on a project for hours, creates slides/sheets/docs/web apps, runs scheduled tasks, and uses desktop computer-use with a built-in browser. Powered by GPT-5.6. Rolling out on web/mobile starting with Pro, Enterprise, and Edu (Plus/Business next); the desktop app is available globally on Mac and Windows for all plans, including Free.
Sites for ChatGPT - ๐ June 2026. A Codex-powered ChatGPT feature that transforms plans and analyses into interactive, sharable websites and lightweight apps. In public beta as of the July 9, 2026 GPT-5.6 / ChatGPT Work launch.
Codex Business Plugins - ๐ June 2026. Enterprise enhancements bringing sales, data analytics, and creative production plugins directly to Codex.
GPT-Rosalind - ๐ June 3, 2026. Major update to OpenAI's life-sciences frontier model โ stronger drug discovery, genomics, quantitative biology, and wet-lab troubleshooting (โ31% fewer tokens than GPT-5.5 on long-horizon genomics analyses). Research preview opened to eligible organizations worldwide; Novo Nordisk joins earlier partners Amgen, Moderna, the Allen Institute, and Thermo Fisher.
GPT-5.5 - ๐ Released April 23, 2026 (codename "Spud"). OpenAI's new frontier model for agentic tasks: coding, online research, data analysis, autonomous tool navigation. Significant gains in reasoning, consistency, and long-horizon task handling. Available on ChatGPT Plus / Pro / Business / Enterprise.
GPT-5.5 Pro - ๐ April 23, 2026. Parallel test-time compute variant for higher-accuracy cognitive tasks. Pro / Business / Enterprise tiers.
GPT-5.5 Instant - ๐ May 5, 2026. New ChatGPT default model. Efficiency-first upgrade with ~50% lower hallucination rate on high-stakes prompts; available on free tier.
GPT-5.5-Cyber - ๐ April 30, 2026. Cybersecurity-specialized variant of GPT-5.5, rolled out via OpenAI's Trusted Access for Cyber (TAC) program to vetted defenders, government, critical infrastructure operators, and security vendors. Not available to the general public.
OpenAI Daybreak - ๐ May 12, 2026. Cyber-defense platform bundling GPT-5.5 + GPT-5.5-Cyber + Trusted-Access-for-Cyber for AI-powered vulnerability detection and patch validation; preview access extended to EU governments and security vendors.
GPT-Realtime-2 - ๐ May 8, 2026. GPT-5-class reasoning brought to the Realtime API, 128K context, parallel tool calls with audio feedback, adjustable reasoning effort.
GPT-Realtime-Translate - ๐ May 8, 2026. Live speech-to-speech translation across 70+ input languages and 13 output languages.
GPT-Realtime-Whisper - ๐ May 8, 2026. Streaming low-latency speech-to-text companion to GPT-Realtime-2.
OpenAI Deployment Company (DeployCo) - ๐ May 11, 2026. New OpenAI-majority-owned services entity for enterprise AI rollout. Backed by $4B+ from TPG / Advent / Bain Capital / Brookfield / Goldman Sachs / SoftBank and consulting partners Bain & Company, Capgemini, McKinsey. Built around Forward Deployed Engineers; absorbs the Tomoro AI consulting acquisition (~150 engineers).
Codex on Mobile - ๐ May 14, 2026. ChatGPT iOS/Android can now remote-control the Codex desktop app โ review outputs, approve actions, switch models, and kick off new tasks from the phone while the live session runs on Mac (Windows next). Rolling out as preview to Free, Plus and Go users.
OpenAI โ Malta partnership - ๐ May 16, 2026. First country-wide deal: every Maltese citizen / resident aged 14+ gets a free 1-year ChatGPT Plus subscription after completing a 2-hour AI literacy course built by the University of Malta. Part of the "OpenAI for Countries" initiative; phased rollout starting May 2026.
OpenAI โ Dell Codex partnership - ๐ May 18, 2026. Brings Codex to hybrid and on-premises enterprise environments via Dell Technologies infrastructure โ first major Codex distribution channel outside the public cloud, targeted at regulated industries needing data-residency control.
ChatGPT Safety Updates โ sensitive-conversation tracking - ๐ May 18, 2026. ChatGPT's safety systems updated to detect and track subtle escalation cues across long sessions for acute risks (suicide / self-harm / harm to others), with cross-session state retention.
OpenAI Guaranteed Capacity (Compute Annual Pass) - ๐ May 19, 2026. Long-term compute reservation product for enterprise AI products / agents / workflows. 1, 2, or 3-year terms; longer terms unlock larger discounts. OpenAI's structural response to the Anthropic "Priority Tier" model.
OpenAI โ Google SynthID + C2PA content provenance - ๐ May 19, 2026. OpenAI partners with Google to add durable cross-platform SynthID watermarking to ChatGPT/Sora images, joins C2PA, and previews a public "is-this-image-from-OpenAI" verifier. First major frontier-lab interop on watermarking.
GPT-5.4 - Released March 2026. Frontier model with 1M-token context, advanced coding, computer use, tool search. BenchLM 94, SWE-bench Verified 77.2%, OSWorld 75% (beats human).
GPT-5.4 Pro - Higher-accuracy variant of GPT-5.4. BenchLM 92.
GPT-5.3 - Early 2026. Includes GPT-5.3 Instant (conversations) and GPT-5.3-Codex (coding).
GPT-5.2 - Released Dec 2025. State-of-the-art reasoning, long-context understanding, and vision.
GPT-5 - Launched August 2025. The default model in ChatGPT, replacing GPT-4o. Multimodal with variants: gpt-5, gpt-5-mini, gpt-5-nano.
GPT-4o - Omni model with native text, vision, and audio. Retired from ChatGPT Feb 2026 but still available via API.
GPT-4.5 - ๐ฆ Retired from ChatGPT late June 2026 (API access continues; conversations auto-migrated to GPT-5.5). Released Feb 2025 as a research preview โ the last GPT-4-family model in ChatGPT. o3 retiring from ChatGPT Aug 26, 2026.
o3 / o4-mini - Reasoning models with chain-of-thought for complex problem solving. Released April 2025. o3 leaves ChatGPT on Aug 26, 2026; the
o3-2025-04-16ando3-pro-2025-06-10API snapshots are removed on Dec 11, 2026, withgpt-5.6-solnamed as the replacement (deprecations).Codex CLI - Open-source terminal-based coding agent powered by OpenAI models.
Anthropic
- Claude Opus 5 - ๐ July 24, 2026. Anthropic's fifth-generation flagship โ nears Fable 5 performance at a significantly lower price ($5/$25 per million input/output tokens). 1M-token context window, 128K output tokens. Now the default model on Claude Max. API:
claude-opus-5. Available on Anthropic API, Amazon Bedrock, and Google Cloud Vertex AI. - Claude Fable 5 (Global Reinstatement) - ๐ July 1, 2026. After US Commerce Department export controls were lifted on June 30, Anthropic reinstated global access to Fable 5 across Claude.ai, the Claude Platform, Claude Code, and Claude Cowork. A new safety classifier blocking the Amazon-discovered jailbreak was deployed (blocks the reported behavior in >99% of cases). Pro/Max/Team and select Enterprise plans got Fable 5 included for up to 50% of weekly usage through July 7, then via usage credits; cloud re-enablement on AWS, Google Cloud, and Microsoft Foundry to follow. Mythos 5 remains restricted to vetted US entities.
- Claude Sonnet 5 - ๐ June 30, 2026. The most agentic Sonnet yet โ planning, browser/terminal tool use, and autonomous operation at a level that recently required Opus-class models. Performance approaches Opus 4.8 on agentic search (BrowseComp) and computer use (OSWorld-Verified) at higher effort settings, with a much wider cost-performance range than Sonnet 4.6. Now the default model for Claude.ai Free/Pro; also on Max/Team/Enterprise, Claude Code, and the API as
claude-sonnet-5. Introductory pricing $2/$10 per million input/output tokens through August 31, 2026 (then $3/$15). Anthropic reports a lower rate of undesirable behaviors than Sonnet 4.6. - Claude Fable 5 - ๐ June 9, 2026. Anthropic's first publicly available Mythos-class model โ a capability tier above Opus. Surpasses Opus 4.8 across software engineering, knowledge work, vision, and scientific research benchmarks. Ships with built-in safeguards (sensitive cyber/bio queries may be rerouted to Opus 4.8). $10 / $50 per million in/out tokens. Available via Anthropic API, Amazon Bedrock, and Google Cloud Vertex AI. โ ๏ธ Access suspended June 12, 2026 โ a US government export-control directive ordered Anthropic to disable Fable 5 and Mythos 5 for all customers pending security review. โ Export controls lifted June 30, 2026; access restored July 1 with a new cybersecurity classifier โ see entry above (statement).
- Claude Mythos 5 - ๐ June 9, 2026. The same underlying Mythos-class model as Fable 5 with fewer restrictions, deployed only to vetted partners (cybersecurity firms, infrastructure providers) through Project Glasswing in collaboration with the US government. Successor to the April Claude Mythos Preview. โ ๏ธ Suspended June 12, 2026 alongside Fable 5 under a US export-control directive. โ Partially reinstated June 26, 2026 โ US Commerce Secretary Lutnick restored access to 100+ approved US companies and federal agencies; broader reinstatement ongoing (statement).
- Claude Opus 4.8 - ๐ May 28, 2026. Major Opus refresh: codebase-scale migrations, sharper agentic judgment, dynamic workflows research preview with hundreds of parallel sub-agents in a single session, manual effort-control panel, 3ร cheaper Fast mode at the same $5 / $25 per million in/out. Available on Anthropic native + Amazon Bedrock + AWS Claude Platform + Google Cloud + Microsoft Foundry. Teases an upcoming Mythos-class model series for limited orgs.
- Claude Opus 4.7 - ๐ Released April 16, 2026. Advanced software engineering (SWE-bench Verified 87.6%), enhanced vision, proactive code verification. Supports
/think xhighreasoning effort. 1M-token context. - Claude Opus 4.6 - Released Feb 2026. 1M-token context, 14.5-hour task horizon. Leads Arena chat leaderboard.
- Claude Sonnet 4.6 - Released Feb 2026. Frontier coding and agentic performance, 1M token context window.
- Claude Mythos Preview - ๐ April 2026 gated research preview. BenchLM 99 (top of leaderboard), SWE-bench Verified 93.9%. Limited to Project Glasswing partners.
- Claude Opus 4 - Released May 2025. Advanced reasoning and complex task execution.
- Claude Sonnet 4 - Released May 2025. Balanced performance and cost for a wide range of tasks.
- Claude Code - Agentic coding tool operating directly in your terminal. Powered by Opus 4.7 with
/think xhighsupport. July 2026: desktop app gains a built-in browser enabling live website interaction (scraping, debugging, live-page inspection); Fable 5 model available since July 1. - Claude Security - ๐ May 1, 2026. Public beta. Enterprise security tool powered by Opus 4.7 โ scans entire codebases for vulnerabilities and generates targeted patches with confidence rating, severity, reproduction steps, and recommended fixes. Available to Enterprise customers via claude.ai/security.
- Claude Finance Agents - ๐ May 5, 2026. Ten Opus-4.7-powered specialised agents for pitchbook authoring, KYC, month-end close, deal screening, etc. Deployable as Claude Cowork plugins, Claude Code skills, or Managed-Agents cookbooks.
- Claude Finance JV - ๐ May 4, 2026. $1.5B Claude deployment joint venture with Goldman Sachs and Blackstone embedding Anthropic engineers in mid-market Wall Street firms.
- Claude Add-ins / Dreaming / Outcomes / Multi-agent orchestration - ๐ May 8, 2026 (Code with Claude 2026). Anthropic introduces Add-ins, scheduled memory review between sessions ("Dreaming"), rubric-driven "Outcomes", and a lead-agent + sub-agent orchestration model with shared filesystem and auditable trace.
- Anthropic โ SpaceX Colossus 1 - ๐ May 6, 2026. Anthropic takes all available capacity at SpaceX's Colossus 1 Memphis datacenter (>220K NVIDIA H100/H200/GB200 GPUs, 300+ MW) for Claude Opus inference. Doubles Claude Code 5-hour rate limits on Pro/Max/Team/Enterprise; also lifts peak-hour limits.
- Anthropic โ AMD (up to 2 GW of Instinct MI450) - ๐ July 22, 2026. Anthropic will deploy up to 2 gigawatts of AMD Instinct MI450 Series (MI455X) GPUs in AMD Helios rack-scale systems with EPYC "Venice" CPUs, Pensando networking and ROCm; the first gigawatt begins in H1 2027. AMD has committed a strategic equity investment of up to $5 billion in Anthropic, plus a multi-year engineering collaboration. Builds on Anthropic's existing MI355X usage โ a deliberate hardware-diversification move alongside its TPU, Trainium and SpaceX Colossus capacity.
- Anthropic's position on open-weights models - ๐ July 27, 2026. Dario Amodei responds to reports that US officials are weighing a ban on Chinese open-weights models: "Anthropic has never advocated for a ban on open-weights models." He calls non-dangerous open weights "a public good" and instead backs chip export controls plus a smuggling crackdown, deterrence of industrial-scale distillation, and mandatory pre-release safety testing for all sufficiently capable models, open and closed. Useful primary source for anyone tracking the 2026 open-vs-closed policy fight.
- Claude for Legal - ๐ May 12, 2026. New legal stack on top of Claude Cowork: 20+ MCP connectors (iManage, NetDocuments, DocuSign, Ironclad, LexisNexis, Westlaw, Harvey, Everlaw, Relativity, CourtListenerโฆ) + 12 practice-area plugins (commercial, employment, privacy, product, corporate, AI governance, litigation associate, law-student bar-exam). Microsoft Word / Outlook / Excel / PowerPoint orchestration built in.
- Claude for Small Business - ๐ May 13, 2026. Small-business toggle inside Claude Cowork โ 15 pre-built agentic workflows across finance / ops / sales / marketing / HR / customer service, native connectors for QuickBooks, PayPal, HubSpot, Canva, DocuSign, Google Workspace, Microsoft 365. Bundled with a free PayPal-backed "AI Fluency for Small Business" course and a 10-city US workshop tour kicking off in Chicago.
- Anthropic โ Gates Foundation $200M - ๐ May 14, 2026. 4-year, $200M partnership pairing grants + Claude usage credits + Anthropic engineers on global-health, life-sciences, education, and agriculture programs. All tools produced under the program will be freely available; first focus areas include vaccine R&D for polio / HPV / preeclampsia and agriculture-specific Claude extensions.
- Anthropic โ PwC strategic alliance expansion - ๐ May 14, 2026. PwC commits to global rollout of Claude Code + Claude Cowork, certifies 30,000 PwC professionals, and stands up a joint "Agentic Enterprise" Center of Excellence โ focused on agentic build, AI-native deals, and finance / supply-chain / HR reinvention.
- Anthropic โ Financial Stability Board briefing (Claude Mythos) - ๐ May 18, 2026. Anthropic briefs the global FSB on Claude Mythos cyber-flaw discovery capabilities โ first time a frontier lab briefs a G20-level financial-stability regulator on a frontier model's offensive-security implications.
- Code with Claude 2026 sessions on YouTube - ๐ May 18, 2026 (sessions published). Full developer-conference recordings (May 6 event) go public: Claude Code roadmap, Claude Developer Platform updates, Managed Agents dreaming + multi-agent orchestration, and partner deployments.
- Widening the conversation on frontier AI - ๐ May 19, 2026. Anthropic publishes its framework for engaging diverse traditions (religious, philosophical, indigenous) in frontier-AI safety dialogue. Companion to ongoing public-engagement work.
- Bristol Myers Squibb โ Anthropic Claude Enterprise - ๐ May 20, 2026. BMS adopts Claude Enterprise as its shared intelligence platform for 30,000+ employees globally, embedding agentic Claude into drug-discovery / development / delivery workflows. First top-5 pharma enterprise-wide Claude deployment.
Google DeepMind
Gemini 3.6 Flash - ๐ โก July 21, 2026. Google's newest Flash tier โ stronger on complex agentic and multimodal tasks while using fewer tokens, at a lower price point than 3.5 Flash. API id
gemini-3.6-flash. Documented in the official Gemini API cookbook alongside thinking-mode guides.Gemini 3.5 Flash-Lite - ๐ July 21, 2026. The fastest, lowest-cost model in the 3.5 family; outperforms prior Flash-Lite generations for high-throughput execution. API id
gemini-3.5-flash-lite. Now the cheapest Gemini tier, superseding 3.1 Flash-Lite for new builds.Gemini 3.1 Pro (preview) - Google's most capable Gemini as of late July 2026, served as
gemini-3.1-pro-preview. GPQA Diamond 94.3% (world-record at launch), ARC-AGI-2 77.1%, BenchLM 94. โ ๏ธ Still carries the-previewsuffix and has no free tier.Gemini 3.5 Pro - โ ๏ธ Delayed โ not yet released as of July 17, 2026 (limited enterprise preview with partners ongoing). Google's upcoming flagship with a reported 2-million-token context window and Deep Think reasoning mode; substantially improved coding and agentic workflow capabilities. Announced at Google I/O in May 2026 for a June release, then postponed after Google scrapped and rebuilt the base model over disappointing coding performance. Third-party reports targeted July 17, but as of July 16 Google says it is still testing 3.5 Pro (alongside an upgraded Flash model) with no confirmed release date. Competes directly with GPT-5.6 Sol and Claude Fable 5.
Gemma 4 12B - ๐ June 2026. Novel multimodal open model with a unified, encoder-free architecture processing text, images, and audio in a single pass. Runs locally on 16 GB VRAM.
DiffusionGemma - ๐ June 2026. 26B MoE open model using text-diffusion for up to 4ร faster generation than autoregressive models.
Gemini 3.5 Flash - ๐ May 19, 2026 โ Google I/O 2026. Default model powering the Gemini app and Google Search AI Mode. Marketed as ~4ร faster than other frontier models in output tokens/sec while outperforming Gemini 3.1 Pro on key benchmarks. Gemini 3.5 Pro was slated for June 2026 but has been delayed (see above).
Gemini Omni / Omni Flash - ๐ May 19, 2026 โ Google I/O 2026. New Google DeepMind multimodal world-model family aimed at AGI. Omni Flash, the first shipped variant, can take any input modality and generate any output (starting with video; image and text generation following). Direct lineage to Gemini Robotics / Genie line of work.
Gemini 3.1 Pro - Released Feb 2026. BenchLM 94, GPQA Diamond 94.3% (world-record), ARC AGI2 77.1%.
$2/1M tokensflagship.Gemini 3.1 Flash Live - ๐ April 2026. Real-time multimodal streaming for voice assistants and interactive agents. Low latency, long context.
Gemini 3.1 Flash-Lite (GA) - ๐ May 8, 2026. Generally available on Gemini API / AI Studio / Vertex AI. Fastest and most cost-efficient model in the Gemini 3 family โ built for low-latency code completion, real-time UX, and agentic developer tools; matches Gemini 2.5 Flash quality at significantly lower cost.
Gemini Omni Flash โ voice-controlled video editing rollout - ๐ May 28, 2026. Omni Flash starts rolling out to consumers via the Gemini app, Google Flow, and YouTube Shorts as the editing engine โ conversational cinematic zooms / background swaps / weather edits driven by text, voice, image, or audio prompts; no traditional NLE required.
Gemini Spark (24/7 personal AI agent) - ๐ May 19, 2026 โ Google I/O 2026. Cloud-resident personal AI agent that runs 24/7 on user intent, integrates Gmail / Chat first, then ~30+ third-party tools via MCP (Adobe / Dropbox / Uber). Available to Google AI Ultra subscribers in the US within the I/O week.
Google AI Ultra ($100/month tier) - ๐ May 19, 2026 โ Google I/O 2026. New top consumer subscription targeted at developers / creators / power users. Gates Gemini Spark, highest Gemini 3.5 quotas, and the upcoming Gemini 3.5 Pro.
Gemini 3.1 Flash / Flash Lite - Fast, cost-efficient models for high-throughput applications.
Gemini 4 (Open) - ๐ Released April 2026. Open model family: 2B / 4B / 26B / 31B variants. Strong science reasoning and document understanding, local deployment ready.
Gemini 2.5 Pro / Flash - GA June 2025. Thinking model with 1M context.
Gemma 4 31B - ๐ April 2026. GPQA Diamond 84.3%. Strong open-weight alternative for on-device reasoning.
Gemma 3 - Previous open model family for on-device and research use.
Gemini Robotics ER-1.6 - ๐ April 14, 2026. Upgraded robotics AI with improved spatial and physical reasoning. Partnership with Agile Robotics for real-world deployment.
Meta
- Muse Image - ๐ July 7, 2026. Meta Superintelligence Labs' most advanced image generation model to date โ an "agentic" image model that performs intermediate reasoning steps (web search, code execution, self-refinement) before producing high-quality visuals. Integrated into the Meta AI app, Instagram Stories (US), and WhatsApp in limited countries (Facebook coming soon). Note: a controversial feature allowing images from other users' public Instagram profiles was added then removed on July 10 after feedback.
- Muse Spark 1.1 - ๐ July 9, 2026. Multimodal reasoning model designed for agentic tasks from Meta Superintelligence Labs โ available through a new public preview of the Meta Model API. Marks a strategic shift toward proprietary revenue-focused models alongside Meta's open-source Llama line.
- Muse Video - ๐ July 7, 2026 (preview). Video generation model from Meta Superintelligence Labs, built on the same foundational technology as Muse Image; ranks #3 on Arena for text-to-video. Previewed alongside the Muse Image launch โ "coming soon to creators and Meta AI."
- Llama 5 โ โ Does not exist. Removed from this list on 2026-07-30 after verification. A "Llama 5, 600B+, April 8 2026" entry circulated widely in AI-news aggregators and LLM search summaries, and was previously listed here. It does not hold up: the
meta-llamaHugging Face organisation contains no Llama-5 weights of any kind (newest Llama-family upload is Llama-4-Maverick, May 2025), and Wikipedia's Llama article states "the latest version is Llama 4, released in April 2025" and that Muse Spark replaced the Llama line in April 2026. Treat any "Llama 5" claim as unverified until Meta publishes weights or a newsroom post. See Muse Spark above for what actually shipped. - Muse Spark - ๐ April 9, 2026. First model from Meta Superintelligence Labs (MSL). Natively multimodal reasoning model powering Meta AI app, smart glasses, and features across Facebook / Instagram / WhatsApp / Messenger.
- Llama 4 Scout - 109B total params (17B active), MoE with 16 experts, 10M token context window, multimodal. Runs on single H100.
- Llama 4 Maverick - 400B total params (17B active), 128 experts, 1M context. Outperforms GPT-4o on multimodal benchmarks.
- Llama 4 Behemoth - 2T parameters (288B active). In training โ Meta's frontier model rivaling top closed-source models.
- Llama 3.3 70B - Strong instruction following and reasoning, open-weight under Llama Community License.
Sakana AI
- Sakana RL Conductor - ๐ Paper April 27, 2026; Fugu beta late-April / early-May 2026. 7B RL-trained orchestrator (built on Qwen2.5-7B) that routes subtasks between GPT-5, Claude Sonnet 4, Gemini 2.5 Pro, etc. SOTA on LiveCodeBench (83.9%) and GPQA-Diamond (87.5%) at ~1.8K tokens/query โ roughly 6ร cheaper than other multi-agent ensembles.
- Sakana Fugu - ๐ Beta April 24-25, 2026. Commercial multi-agent orchestration service productising the RL Conductor research. OpenAI-compatible API with two tiers: Fugu Mini (low-latency) and Fugu Ultra (max performance); strong reported results on SWE-Pro, GPQA-D and ALE-Bench.
Zyphra
- ZAYA1-8B - ๐ May 6, 2026. MoE reasoning model (<1B active) trained end-to-end on AMD Instinct MI300X clusters. Apache 2.0 weights on Hugging Face + serverless endpoint on Zyphra Cloud; aimed at math, code, and dense reasoning per active parameter.
- ZAYA1-8B-Diffusion-Preview - ๐ May 14, 2026. First MoE diffusion language model converted from an autoregressive LLM and the first diffusion LM trained on AMD GPUs. Generates 16 tokens per step, achieving up to 7.7ร inference speedup vs the autoregressive base. Built with Zyphra's TiDAR recipe + CCA attention.
Thinking Machines Lab
- Inkling - ๐ July 15, 2026. Founded by Mira Murati (former OpenAI CTO). 975B MoE parameters (41B active), pretrained on 45T tokens, 1M-token context window. Natively multimodal โ text, image, audio, and video in a single model. Apache 2.0 open weights on Hugging Face. Also ships Inkling-Small (12B active parameters). Available via the Thinking Machines API and Hugging Face Inference.
- Inkling-Small - ๐ โก July 30, 2026 (weights released). Compact variant of Inkling โ 276B total / 12B active, same native multimodal architecture (text/image/audio), 1M-token context, Apache 2.0. Scores 31.6% on HLE text benchmark โ slightly outperforming the larger 975B Inkling (29.7%) on that metric, validating the efficiency-first design. Available via Thinking Machines API and Hugging Face.
Mistral AI
- Mistral Large 3 - 675B total / 41B active parameters, MoE, 256K context. Flagship open-weight multimodal model. Released Dec 2025.
- Mistral Medium 3.1 - Frontier-class dense model for enterprise. Multimodal, 128K context, 80+ coding languages. Released Aug 2025.
- Mistral Small 4 - ๐ Released March 2026. 119B total / 6B active. Hybrid model combining reasoning, multimodal, and coding strengths.
- Magistral 1.2 - ๐ 2026 reasoning family challenging o3/o4-mini. Transparent and multilingual reasoning.
- Devstral 2 - ๐ 2026 agentic coding model. Best open-source model for coding agents.
- Codestral - 22B code generation model, 80+ programming languages, 32K context. Released May 2024.
- Pixtral Large - 124B multimodal model with 1B vision encoder, 128K context, processes 30+ high-res images.
- Ministral 3B/8B/14B - Compact models optimized for edge deployment and efficiency.
- Mistral Forge - ๐ March 2026 platform for training custom LLMs on proprietary data.
- Mistral Medium 3.5 - ๐ April 28, 2026. Dense 128B open-weight model, 256K context, Modified MIT license. Unifies instruction-following, reasoning, and coding.
- Leanstral 1.5 - ๐ July 2, 2026. Formal-verification model for proof engineering in Lean 4 โ 119B total / 6B active parameters, Apache 2.0, weights on Hugging Face plus a free API endpoint. Scores 100% on miniF2F, solves 587/672 PutnamBench problems, and discovered 5 previously unreported bugs across 57 real-world repositories.
- Robostral Navigate - ๐ July 8, 2026. Mistral's first robotics model โ an 8B embodied-navigation model that steers wheeled, legged, and flying robots through offices, homes, and outdoor spaces from natural-language instructions using only a single RGB camera (76.6% success rate on unseen validation). Trained fully in-house on ~400K simulated trajectories.
- Voxtral TTS - ๐ March 26, 2026. 4B-parameter open-weight TTS built on Ministral 3B; multilingual, optimised for voice agents.
DeepSeek
- DeepSeek-V4-Pro - ๐ April 24, 2026 (preview); production launch mid-July 2026. 1.6T total / 49B active MoE, 1M-token context. MIT license. Leadership in agent capabilities, world knowledge, reasoning; tops open-source benchmarks. Official pricing is flat (no peak/off-peak tiering): $0.003625 cache-hit / $0.435 cache-miss input, $0.87 output per 1M tokens, 384K max output, 500-request concurrency.
deepseek-v4-pro/deepseek-v4-flashare the production API models. - DeepSeek-V4-Flash - ๐ April 24, 2026. 284B total / 13B active MoE, 1M context. MIT. Cost-efficient tier โ $0.0028 cache-hit / $0.14 cache-miss input, $0.28 output per 1M tokens, 384K max output, 2,500-request concurrency (pricing).
- DeepSeek-V4-Flash-0731 - ๐ โก July 31, 2026. Updated Flash checkpoint with enhanced agentic capabilities โ same 284B/13B-active MoE architecture, same pricing/API model ID, but outperforms V4-Pro (Preview) on agent task benchmarks. Open weights on Hugging Face under MIT license. Drop-in replacement for
deepseek-v4-flashAPI users. - DeepSeek Agent Harness team - ๐ May 19, 2026. DeepSeek hires a former Jane Street engineer to lead a new "AI harness" team building the deterministic scaffolding that turns DeepSeek V4 into autonomous, revenue-generating agents โ first major signal DeepSeek is moving past raw-model R&D into agentic productisation.
- DeepSeek-V3.2 - Released Dec 2025. Advanced MoE architecture with 671B total parameters. V3.2 Speciale variant for enhanced reasoning. โ ๏ธ API model IDs deepseek-chat / deepseek-reasoner (V3.2-era) deprecated effective July 24, 2026 โ superseded by V4-Flash modes.
- DeepSeek-R2 - ๐งช Unreleased/rumored. No official announcement, model card, or API ID exists as of mid-July 2026; reasoning is served via V4's Thinking mode.
- DeepSeek-R1 - Reasoning-focused model with chain-of-thought capabilities. Released Jan 2025.
- DeepSeek-Coder-V2 - Code generation model competitive with GPT-4 on coding benchmarks.
Alibaba (Qwen)
- Qwen3.7-Max - ๐ May 20, 2026 โ Alibaba Cloud Summit Hangzhou. New Qwen flagship purpose-built as the foundation for AI agents: agentic coding, complex reasoning, and long-horizon multi-step missions with sustained decision-making. Released alongside a full-stack AI infrastructure upgrade and new T-Head Zhenwu M890 AI accelerator chip. Worldwide developer/enterprise availability rolling.
- Qwen3.7-Max-Preview / Qwen3.7-Plus-Preview - ๐ May 18, 2026. Preview ladder before the Hangzhou unveil. Ranked the highest of any Chinese model on LM Arena in both text and vision; sustained 1M-context evaluations.
- Qwen3.6-27B - ๐ April 22, 2026. Dense 27B multimodal. Open-sourced. Focus: agentic coding + thinking-context preservation.
- Qwen3.6-Max-Preview - ๐ April 18, 2026. Proprietary frontier preview. High coding/reasoning performance, 1M context window. Top-tier among Chinese models on coding benchmarks.
- Qwen3.6-35B-A3B - ๐ April 15, 2026. MoE, 35B total / 3B active. Apache 2.0. Stability and real-world utility improvements.
- Qwen3.6-Plus - ๐ April 2, 2026. Proprietary flagship. High value-per-token general model. Strong long-context, tool-calling, agentic behavior.
- HappyHorse 1.1 - ๐ June 23, 2026. Alibaba's video-generation model (T2V/I2V/S2V, up to 15s 1080p with synced audio, strong multi-shot character consistency). HappyHorse 1.0 entered limited beta April 28, 2026 after launching anonymously and topping video leaderboards.
- Qwen3.5 Max Pro - April 2026. High-performance flagship. Enhanced coding and math reasoning, long context.
- Qwen3.5 Omni Plus - April 2026. Proprietary full-modal foundation model unifying text and image input.
- Qwen3-Max-Thinking - Alibaba's strongest thinking model. 1T+ parameters, enhanced agentic capabilities.
- Qwen3.5-Omni - March 2026. Fully omni-modal: language, vision, sound, motion. Speech recognition in 113 languages, 256K context.
- Qwen3-Coder-Next - Feb 2026. Open-weight coding agent model, MoE 80B total / 3B active.
- Qwen3 235B-A22B - MoE with dual-mode reasoning. Strong math, code, and commonsense reasoning.
- Qwen2.5 Coder 32B - Top open-source coding model.
xAI (Grok)
- Grok 4.5 - ๐ July 8, 2026. xAI's latest flagship โ optimised for coding and agentic tasks through joint training with Cursor using real developer interaction data. Features a 500K-token context window, function calling, structured outputs, web/X search, code execution, document search, and context compaction. Available in xAI API, Grok Build, and as the default model in Cursor. Priced at $2/$6 per million in/out tokens. EU availability expected later in July 2026.
- Grok 4.3 GA - ๐ May 2026. Grok 4.3 reached general availability on Microsoft Foundry and OCI Generative AI; xAI's flagship for agentic workloads with improved tool-calling and long-horizon reasoning.
- Grok 4.3 Beta - ๐ April 2026. Latest iteration with improved reasoning and coding benchmarks. See
2026.4benchmark snapshot. - Grok 4.20 - Feb 2026. Multi-agent system (4 standard + 16 specialized agents in Heavy mode), 2M token context.
- Grok 4 / 4 Heavy - Released July 2025. xAI's frontier model of the Grok 4 generation.
- Grok 3 / 3 Mini - Feb 2025. First reasoning models with "Think Mode".
Microsoft (MAI)
- Microsoft MAI-Code-1-Flash - ๐ Build 2026 (June 2, 2026). Microsoft's first major in-house foundation model built entirely without OpenAI technology. 5B-parameter coding model with adaptive thinking, rolling out in GitHub Copilot. Outperforms Claude Haiku 4.5 across four core coding benchmarks (16-point lead on SWE-Bench Pro: 51.2% vs 35.2%); solves harder tasks with up to 60% fewer tokens on SWE-Bench Verified.
- Microsoft MAI-Thinking-1 - ๐ Build 2026 (June 2, 2026). Microsoft's first in-house reasoning model, trained from scratch without OpenAI data. Companion to MAI-Code-1-Flash; signals Microsoft's foundation-model independence push.
Microsoft (Phi)
- Phi-4-reasoning-vision-15B - ๐ Released March 2026. 15B multimodal model with selective chain-of-thought reasoning. Edge-deployable.
- Phi-4 - 14B parameter SLM with reasoning rivaling much larger models. Open-source under MIT License.
- Phi-4-mini - 3.8B parameter dense model. 128K context. Excels in reasoning, math, coding, and function-calling.
- Phi-4-multimodal - 5.6B parameter. First multimodal Phi model โ integrates speech, vision, and text in unified architecture.
Cohere
- Command A+ - ๐ May 20, 2026. 218B total / 25B active MoE, Apache 2.0 open weights (Hugging Face). 128K input / 64K output. Multimodal, 48 languages, agentic tool use; runs on 2ร H100 or 1 Blackwell GPU.
- Command A - Released March 13, 2025. 111B open-weights model, 256K context. Agentic, multilingual, and coding focused.
- Command R+ - Enterprise RAG model, 128K context, multilingual (10 languages), grounded generation with citations.
- Command R - Cost-efficient model for retrieval-augmented generation and enterprise workloads.
Baidu (ERNIE / ๆๅฟ)
- ERNIE 5.1 - ๐ May 8, 2026. ~1/3 the total and ~1/2 the active parameters of ERNIE 5.0 at ~6% of comparable pre-training cost; #1 Chinese model / #4 global on LMArena Search (1,223).
- ERNIE 5.0 - Released November 13, 2025 (Baidu World). 2.4T-parameter omni-modal MoE (activates <3% per query).
- ERNIE 4.5 - Multimodal predecessor released 2025. Strong reasoning and Chinese language capabilities.
Zhipu AI / Z.ai (GLM)
- GLM-5.2 - ๐ June 13, 2026. Coding-first 744B-MoE flagship with a 1M-token context window (~5ร GLM-5.1) and up to 131K output tokens. Live across all GLM Coding Plan tiers; MIT open weights + standalone API rolling out the launch week. Works out of the box with Claude Code, Cline, OpenCode, Roo Code, Goose, and OpenClaw. (No benchmark numbers published at launch.)
- GLM-5.1 - ๐ April 8, 2026. 744B MoE / 40B active, 200K context. MIT license. Tops SWE-Bench Pro.
- ZCode - ๐ ๐จ๐ณ July 2, 2026. Zhipu's agent harness for GLM-5.2 โ turns the model into an autonomous coding agent, squarely targeting Claude Code; launch promos include +50% quota for Coding Plan subscribers and 5M free tokens for new users.
- GLM-5 Reasoning - ๐ April 2026. BenchLM 85 โ top open-source score. SWE-Bench Pro surpasses GPT-5.4 and Claude Opus 4.6.
- GLM-5V-Turbo - ๐ April 2026. Native multimodal agent โ vision, video clips, text inputs. Cost-performance balanced.
- GLM-5 - Released Feb 2026. 744B parameters, advanced agentic intelligence. MIT license.
- GLM-4.7 - Released late 2025. Matches Claude Opus 4 on SWE-Bench.
MiniMax
- MiniMax M3 - ๐ ๐จ๐ณ June 1, 2026. Open-weight flagship with MiniMax Sparse Attention โ ~1/20 the compute cost at 1M tokens; frontier coding capabilities. Weights at
MiniMaxAI/MiniMax-M3(82 files, ungated, ~155K downloads). โ ๏ธ Licence corrected 2026-07-30 โ this is not MIT. The model card declareslicense: other/license_name: minimax-community, i.e. a bespoke community licence, so read its terms before commercial use. (Its SWE-bench Pro figure is also best ignored โ see the benchmark caution.) - MiniMax-M2.7 (Open Weights) - ๐ April 2026. 230B-class open-weight flagship. Top-tier performance on coding and Agent tasks.
- MiniMax M2.7 - ๐จ๐ณ ๐ March 2026. Proprietary self-evolving LLM tuned for agent harness construction, memory updates, iterative workflow improvement; major gains on SWE-bench-style tasks.
- MiniMax M2.5 - ๐จ๐ณ February 2026. 230B-parameter cost-efficient flagship for "real-world productivity".
- Hailuo 2.3 / 2.3 Fast - ๐จ๐ณ October 2025. MiniMax's current video flagship โ SOTA physics, character micro-expressions, strong stylization; Hailuo 02 (2025) remains as the I2V-focused variant.
- MiniMax Music 2.6 - ๐จ๐ณ ๐ April 10, 2026. Cover-generation focus with improved low-frequency reproduction; global beta.
- MiniMax-M1-80k - Open-weight hybrid-attention reasoning model. 456B parameters, 1M token context.
- Hailuo AI (Video) - Text/image-to-video generation with AI avatars, voiceovers, and character consistency.
- Kilo Code Integration - MiniMax models are heavily featured in Kilo Code (open-source AI coding extension at kilo.ai).
Moonshot AI (Kimi)
This HTML preview is truncated for page performance. The canonical Markdown file contains the complete snapshot.
Why MDRSS assigned this score
- Production catalog audit 2026-08-04
- Taxonomy classified from title, annotation, source and Markdown signals
- Agent usefulness evaluated from structure, procedures, examples, evidence and retrieval value
Discussion 0
Sign in to join the discussion.