What is an agent harness?

Snapshot 2026-08-03 23:56:39 UTC · version 1

published
INDEXABLE MARKDOWN SNAPSHOT

Research document

Open canonical .md

Best of Agent Harnesses and Harness Techniques

🏆  Curated list of AI agent harnesses, orchestration frameworks, and harness techniques for reliable agentic systems.

🌐 Browse the searchable site — one page per harness, filter by capability, autonomy & recovery.

🤖 Agents can query this list — an MCP server (recommend, pick_harness, …), llms.txt & JSON, so your agent recommends harnesses too.

🧡 A curated list is only as good as the people who stop mid-scroll to point at what it's missing.
These folks did exactly that — found a gap, wrote it up, and made the list better than one maintainer ever could. Meet the 7 →

What is an agent harness?

A model answers; an agent acts. An agent harness is the runtime that turns one into the other — the model thinks; the harness decides what that thinking is allowed to touch.

Every prior wave of automation was constrained by brittleness: you scripted exact behavior, and when the world deviated, the system broke. Foundation models inverted that problem—they're flexible but directionless, stateless, and disconnected from anything real. The agent harness exists to bridge that gap: it is the orchestration infrastructure that converts a model's per-turn reasoning into sustained, tool-using, error-recovering, goal-directed behavior across time. Architecturally, it plays the role the kernel played in operating systems or the controller played in industrial robotics—mediating between raw capability and a messy environment—but with a critical difference: the "capability" it governs is general-purpose cognition, which means the harness is simultaneously a scheduler, a permission system, a memory manager, and a policy enforcement layer, all under-specified and evolving in real time.

Why harnesses matter

Better models make harnesses more important: more capabilities mean more failure modes, and production needs retry logic, fallbacks, and validation. Harness quality—not just model quality—determines whether agents actually ship. This list ranks projects by relevance to harness concerns (environment, orchestration, lifecycle, guardrails) and by stars/activity.

The landscape at a glance

Every project in the list, plotted by adoption surface area (the simplicity ↔ capability axis) against GitHub stars. Colors are categories; the largest projects in each tier are labeled.

The same projects placed by how much unsupervised rope they're designed to give (autonomy) and what happens when a run dies (recovery). In the tables below, ★ marks headless-ready projects and ✱ marks durable ones. Both charts regenerate from the list data on every refresh.

How to Pick a Harness

Start with the guide, then the head-to-head decision pages — grounded in the same data as the tables below:

Pick by use case

Reader's index: pick by what you want to do, not by category. Tag chips (e.g. mcp · memory) next to each row let you cross-filter by capability — see TAGS.md for the full cross-reference.

For agents

This list is also published in machine-readable form, so coding agents and research agents can recommend harnesses — not just humans browsing GitHub:

  • harnesses.json — every project with category, complexity tier, capability tags, stars, license signal, and a concrete example link, plus the full use-case index.
  • llms.txt — the entire list in one agent-readable file. Point any agent at the raw URL.
  • MCP serverrecommend (one opinionated pick + alternatives + what to avoid, e.g. repos flagged for star manipulation), compare/compare_for (2–4 harnesses side by side — by id or by task — who leads on which axis incl. researched sandboxing/memory/hooks/prompt-optimization ratings, graveyard warnings, the matching decision guide), pick_harness (ranked, with complexity/autonomy/recovery filters), search_harnesses, get_harness, list_categories, plus list_comparisons/get_comparison for the decision guides. Published to PyPI and the official MCP registry as io.github.RyanAlberts/agent-harnesses. One-line install (needs uv):
claude mcp add agent-harnesses -- uvx agent-harnesses-mcp

Or hire a skeleton

Don't just read the list — agents/ ships three agent skeletons: open-source agents that run on the AI subscription you already pay for. Clone the file, customize the instructions, done. All three work against the current week's data and deliver to Slack or Notion when either is connected:

  • harness-scout — describe what you're building; it picks your harness, with evidence and a graveyard check.
  • stack-auditor — flags the harnesses in your codebase that died, and can trace your agent session logs to show how the harness steers your technical decisions.
  • harness-radar — weekly movement briefing: climbers, arrivals, deaths, graduations.
curl -fsSL https://raw.githubusercontent.com/RyanAlberts/best-of-Agent-Harnesses/main/agents/harness-scout.md -o .claude/agents/harness-scout.md

Contents

Guide to rankings

  • Stars — GitHub star count, captured 2026-08-02; tables sort by stars descending.
  • ⚖️ Simplicity ↔ capability — adoption surface, 4 tiers: super simple (a format, one concept) → mostly simple (thin layer) → slightly complex (real SDK) → complex (product suite).
  • Headless-ready — designed for unattended runs, batches, and fleets (the top of the autonomy scale: step-gated → checkpoint-gated → bounded → headless).
  • Durable — persisted execution state survives restarts mid-task (the top of the recovery scale: none → retry → resumable → durable).
  • Open source — ✅ standard OSS license · ⚠️ source-available/restricted · ❓ no or unclear license.
  • 🏷️ Tags — capability chips auto-derived from descriptions; full cross-reference in TAGS.md.
  • 🎯 Examples — one concrete "show me it in action" link per project, not a docs root.

Every project's full autonomy and recovery tier is plotted in the grid above and carried in harnesses.json and llms.txt; scores are editorial, from public docs — maintainer corrections via issue/PR are merged fast.


Progressive disclosure harnesses

Formats, runtimes, and patterns that reveal context, tools, or instructions in layers—index first, details on demand—to control tokens and improve agent focus (the "map, not encyclopedia" principle).

# Project ⭐ Stars Description Open source Simplicity ↔ capability Examples
1 Headroom 64k Compresses tool outputs, logs, files, and RAG chunks with content-aware compressors before they reach the model—claimed 20% fewer tokens for coding agents and 60–95% fewer for JSON, same answers. Ships as a library, HTTP proxy, or MCP server, so it drops in front of whatever harness you already run. mcp · rag mostly simple (compression library/proxy/MCP server) Project README
2 awesome-cursorrules 40.5k Curated .cursorrules and skills that leverage Cursor's index-then-load model; the canonical collection for rules-as-progressive-disclosure in the IDE. ide super simple (content bundle) PyTorch cursorrules
3 agents.md 23.4k Open format for repo-scoped agent briefings; v1.1 adds hierarchical scope and progressive disclosure so agents get a map of what exists, then load only what's relevant. typescript super simple (format only) Self-hosting AGENTS.md
4 context-mode 19.6k Context-window optimization layer that sandboxes tool output before it reaches the model (claimed 98% reduction) and persists session memory across 17 agent platforms via MCP and hooks—progressive disclosure applied to tool results, not just instructions. mcp · memory · sandbox mostly simple (output sandboxing, cross-platform) Project README
5 langgraph-bigtool ✱ 552 Build LangGraph agents with large tool sets; retrieval and on-demand tool loading so agents scale beyond context without stuffing every schema upfront. tool-discovery · python slightly complex (large tool sets) Math-library tool agent
6 MCP-Zero 503 Active tool discovery for autonomous agents: model requests tools by requirement; hierarchical semantic routing over 308 servers / 2,797 tools with ~98% token reduction (APIBank). tool-discovery complex (3k tools, full routing) APIBank experiment
7 ToolGen 183 ICLR 2025: unified tool retrieval and calling via generation; 47k+ tools without context stuffing—retrieval and invocation in one generative step. tool-discovery · python complex (47k+ tools) Full eval pipeline
8 ToolRAG 29 Semantic tool retrieval for LLMs; serves only the tools the user query demands (MCP-compatible), unlimited tool sets with zero context penalty. mcp · tool-discovery mostly simple (query-driven retrieval) MCP server retrieval

Coding agent products (IDEs, CLIs, full suites)

Turnkey coding agents you install and run: IDE extensions, terminal CLIs, Dockerized workspaces. Each entry notes which part is the harness (the agent loop, tool wiring, approval model) versus the UI shell (VS Code extension, TUI, browser client).

# Project ⭐ Stars Description Open source Simplicity ↔ capability Examples
1 opencode ★ 192k Open-source terminal coding agent (formerly sst/opencode; transferred to anomalyco). The harness is a multi-provider tool-call loop (Claude, OpenAI, Gemini, local) with strong plugin and MCP support; the TUI is the shell. 100% OSS, very actively shipped. mcp · provider-agnostic · cli · tui · typescript slightly complex (multi-provider, plugins, MCP) Agent system page
2 Gemini CLI 106k Google's first-party terminal agent for Gemini. The harness is the plugin/MCP tool-call loop; the terminal is the shell—Google's parallel to Claude Code / Codex, not just an API. mcp · cli · typescript slightly complex (official CLI, plugins, MCP) MCP server setup
3 Codex 103k OpenAI's terminal coding agent. The harness is the sandboxed tool-call loop with multi-provider support; the CLI is the shell. Reference implementation for "official CLI that ships code." sandbox · provider-agnostic · cli slightly complex (reference CLI, sandboxed) Sandboxing concept
4 OpenHands ★ 82.9k Dockerized software-engineering agent. The harness is the bash/editor/browser toolset with micro-agents and event-stream session bridging; Docker is the sandbox. Main OSS choice for teams self-hosting autonomous repo work. memory · browser · sandbox · python ⚠️ (multi-license) complex (Docker runtime, multi-surface agent — product suite) Repository microagents
5 pi 82.2k The upstream AI agent toolkit behind this list's oh-my-pi fork: a unified multi-provider LLM API, agent loop, and TUI shell providing the harness that oh-my-pi's Rust rewrite builds on. provider-agnostic · tui · rust slightly complex (multi-provider agent loop, TUI) Project README
6 Open Interpreter 67.5k Lightweight terminal coding agent oriented to open models (DeepSeek, Kimi, Qwen). The harness is a code-execution loop — the model writes code, the harness executes it with confirmation gates; the CLI is the shell. The original "let the LLM run code on my machine" project, reborn for open weights. cli · python mostly simple (lean code-exec loop) Quick start
7 Cline 65.5k VS Code extension whose harness is a plan-then-act loop with per-step human approval and cost transparency; the VS Code integration is the UI shell. Open-source counterweight to Cursor. ide · typescript slightly complex (plan-then-act, approval gates) Plan & Act mode
8 goose ★ 52.1k Block-originated Rust agent, now stewarded by the Linux Foundation's Agentic AI Foundation (aaif-goose/goose). The harness is the MCP/ACP extension model with recipes and provider choice; there's no fixed UI slot—you bolt it into whatever shell you use. mcp · rust slightly complex (extensions, MCP/ACP) Goose recipes guide
9 DeepSeek-Reasonix 28.8k DeepSeek-native terminal coding agent. The harness is engineered around prefix-cache stability for long-running sessions; the TUI is the shell. memory · cli · tui · typescript slightly complex (terminal agent, prefix-cache tuned) Project README
10 vibe-kanban 27.6k Kanban-style fleet manager for running Claude Code, Codex, or any coding agent across many tasks at once. The harness contribution is the task-queue/review layer on top of whichever agent executes; not an agent loop itself. slightly complex (task-fleet manager) Project README
11 crush 27k Charm's terminal coding agent (Charm's fork of the original OpenCode). The harness is the tool-calling loop with session persistence; the Bubble Tea TUI is the shell. memory · cli · tui ⚠️ FSL-1.1-MIT slightly complex (terminal agent, TUI) Crush launch post
12 Kilo Code 26.7k VS Code extension and CLI in the Cline/Roo-Code lineage — a natural pick now that Roo-Code is archived upstream. The harness is an approval-gated autonomous-mode loop with a provider/tool marketplace; the IDE is the shell. mcp · cli · ide · typescript slightly complex (IDE extension + CLI, MCP) Project README
13 qwen-code 26.5k Alibaba's official terminal coding agent, forked from Gemini CLI's agent loop and retuned for Qwen models. The harness is the same sandboxed tool-call loop as its upstream; the terminal is the shell. sandbox · cli · typescript slightly complex (official CLI, Gemini-CLI fork) Project README
14 Symphony ★ 26.4k OpenAI's harness for fanning a task out into many isolated, autonomous coding-agent implementation runs and surfacing the ones that pass, so a team manages outcomes instead of supervising each session. sandbox complex (parallel isolated runs — product suite) Project README
15 Roo Code 24.4k VS Code/Cursor extension in the Cline lineage. The harness is the approval-gated agent with custom modes and a strong MCP story; the IDE is the UI. Popular community fork when you want that workflow without the upstream extension. mcp · workflow · ide · typescript slightly complex (IDE extension, MCP-first) Custom modes guide
16 oh-my-pi 21.2k Terminal coding agent (fork of Pi) that wires the IDE into the harness: hash-anchored edits, a 32-tool loop tuned per-model, LSP rename/references/diagnostics on every write, a real DAP debugger (lldb/dlv/debugpy), long-lived Python + Bun execution kernels that call back into the agent's tools, browser control, and 40+ providers (Claude/OpenAI/Gemini/local). ~55k-line Rust core. browser · provider-agnostic · cli · ide · rust slightly complex (terminal agent, LSP/DAP, multi-provider) LSP wired into edits
17 jcode 15.2k Rust terminal coding agent pitched as the most RAM-efficient harness in its class; MCP support, multi-provider (Claude/OpenAI). mcp · memory · provider-agnostic · cli · rust slightly complex (terminal agent, low-memory) Project README
18 eigent 14.7k Open-source desktop harness positioned as a local, free alternative to Claude Cowork and Codex: multi-agent workspace orchestration in a self-hosted app rather than a hosted product. multi-agent · local complex (desktop multi-agent workspace — product suite) Project README
19 cc-haha 13.9k Local-first desktop workspace harness for Claude Code and other agents: multi-agent sessions, Git worktrees, code diffs, a skill marketplace, and chat-app access (WeChat, Telegram, WhatsApp). memory · multi-agent · typescript complex (desktop workspace, multi-agent — product suite) Project README
20 claw-code-agent 537 Python reimplementation of the Claude Code agent architecture with zero external dependencies; interactive chat, streaming, plugin runtime, nested agent delegation, cost tracking, MCP transport—portable harness without the Rust/TS toolchain. mcp · rust · python · typescript slightly complex (pure Python, plugin runtime) Quick Start guide
21 AgentBox 332 Runs multiple coding agents in parallel, each in its own sandboxed VM, locally or in the cloud, from one command. The harness contribution is the VM-per-agent isolation and fleet fan-out layer; whichever agent runs inside owns the loop. sandbox · typescript slightly complex (VM-per-agent sandbox, parallel fan-out) Parallel agents quick start
22 Proliferate 156 Open-source AI IDE for Claude Code, Codex, OpenCode, and more. The harness contribution is the workspace/session orchestration layer: run multiple coding agents in parallel, locally or in the cloud, with isolated workspaces, reusable workflows, and shared team context. multi-agent · sandbox · ide · typescript complex (multi-agent workspace orchestration — product suite) Product README

Coding harness configs and SDKs

Skill packs, slash-command libraries, meta-prompting frameworks, and official SDKs that give you the harness (the agent loop, planning, memory, hooks) without bundling a specific IDE or CLI shell.

# Project ⭐ Stars Description Open source Simplicity ↔ capability Examples
1 superpowers 265k Performance-oriented harness pack for Claude Code, Codex, OpenCode, Cursor: skills, instincts, memory, security, research-first workflows. Treats harness engineering itself as the performance lever. memory · ide complex (multi-IDE skill stack — product suite) TDD skill
2 Anthropic Skills 166k Anthropic's official Agent Skills repository: SKILL.md-based folders (instructions, scripts, resources) Claude dynamically loads on Claude Code, Claude.ai, and the API. The reference for progressive-disclosure skill packs in 2026. mostly simple (official skills format) docx skill
3 GStack 126k Garry Tan's Claude Code skill stack: 23 slash-command modes (CEO/eng/design review, QA, ship, browse, retro, …) that structure one assistant as a virtual engineering team. Daily driver while running YC. typescript slightly complex (multi-role slash-command harness) /ship SKILL.md
4 addyosmani/agent-skills 81.3k Addy Osmani's production-grade skill pack: 24 engineering skills and 4 specialist agent personas that encode senior-dev workflows (spec through deploy) across 70+ coding agents including Claude Code, Cursor, and Copilot. The harness contribution is the skill/workflow layer, not a new agent loop. workflow · ide mostly simple (skills bundle, cross-agent) Project README
5 awesome-claude-code 51.5k Large community-curated index of Claude Code skills, slash commands, status lines, and plugins—resources for extending the harness, not a harness itself, but the most-followed catalog of the genre. super simple (curated resource index) Project README
6 wshobson/agents 38.4k Cross-harness marketplace of drop-in subagents and skills for Claude Code, Codex CLI, Cursor, OpenCode, and Copilot; specialized, production-ready agent definitions you install rather than hand-write. multi-agent · cli · ide super simple (drop-in agent packs) Agent catalog
7 planning-with-files 25.9k Skill for persistent, file-based planning across long-running coding-agent sessions: crash-proof markdown plans, session recovery after /clear/compaction, and a deterministic completion gate—Manus-style planning as a drop-in harness layer via the Agent Skills standard. memory mostly simple (skill, file-based state) Project README
8 SWE-agent ★ 20k LM-driven harness built for SWE-bench: edit state, command execution, and issue-focused loop—the reference agent stack next to the benchmark itself. memory · evals · python slightly complex (SWE-bench pairing, stateful edits) Default agent config
9 Claude Agent SDK ★ 7.8k Official Anthropic SDK (Python + TypeScript, demos, quickstarts): built-in tools, MCP, long-running coding agents with session bridging. mcp · memory · python · typescript complex (full SDK, session bridging — product suite) Research agent demo
10 get-shit-done 7.6k Goal-backward planning and wave-based execution over fresh context windows; avoids context rot by design. Python/JS meta-prompting for Claude Code, OpenCode, Gemini CLI. cli · python mostly simple (meta-prompting, you own stack) gsd:ship command
11 agents-cli 5.5k Google's official CLI and skill pack that layers agent-creation, evaluation, and deployment skills on top of whatever coding assistant you already run, rather than shipping its own agent loop—the harness as a config/skills add-on, not a new runtime. evals · cli mostly simple (skills/CLI layer, no new runtime) Project README
12 skillhub 4.8k iFlytek's self-hosted registry for publishing, versioning, and governing agent skill packages—the harness config layer treated as an enterprise artifact store rather than a CLI or IDE shell. local · cli · ide mostly simple (skill registry/governance) Project README
13 Meta-Harness 1.4k Reference implementation from the Meta-Harness paper: an academic testbed for harness-engineering research, not a product—useful as a citation-grade baseline rather than something you'd run in production. slightly complex (research reference implementation) Project README
14 RepoMaster ★ 541 Repo-scoped research harness: builds function-call and module-dependency graphs to explore only what's needed; large relative gains on MLE-bench and GitTaskBench with lower token use. workflow · python slightly complex (graph-based exploration) PDF-parse case study
15 AutoHarness 363 Lightweight governance harness: wraps any LLM client in ~2 lines for automated harness engineering—6–14 step pipeline, YAML constitution, risk-pattern matching, session persistence with cost tracking, multi-agent profiles. memory · multi-agent · provider-agnostic · python super simple (2-line wrapper, YAML gov) Full pipeline demo
16 LoopTroop 106 Config layer that chains LLM councils for planning, Ralph loops for iterative refinement, and OpenCode worktrees for shipping. The harness contribution is the council → loop → worktree pipeline; OpenCode underneath executes. typescript mostly simple (config pipeline over OpenCode) Council → loop → worktree pipeline
17 pmstack 8 Claude Code config for AI product managers: CLAUDE.md plus skills for competitive analysis, PRD-from-signal, metric frameworks, stakeholder briefs, and agent eval design. "GStack for PMs." evals super simple (skills bundle, PM-focused) PRD-from-signal skill

Personal agent runtimes

Always-on, self-hosted agents you run as a daemon and talk to from chat apps: gateway runtimes, second brains, and self-improving assistants. The agent as a product you operate, not a library you build with.

# Project ⭐ Stars Description Open source Simplicity ↔ capability Examples
1 OpenClaw ★ 385k Self-hosted, always-on personal agent (formerly Clawdbot/Moltbot): a gateway + event-loop runtime that treats messages, heartbeats, crons, and webhooks as one input queue, persists state to local files, and lives in your chat apps (WhatsApp, Telegram, Slack, Discord). 13,700+ community skills; the fastest-growing repo in GitHub history. typescript · multi-agent complex (always-on runtime, channels, skill ecosystem — product suite) Agent runtime architecture
2 Hermes ★ 224k Nous Research's self-improving agent: a learning loop turns experience into reusable skills, builds a persistent user model across sessions, and checkpoints state to disk with rollback; lean enough for a $5 VPS, driven from chat, and model-agnostic (Nous Portal, OpenRouter, OpenAI, or any endpoint). memory · python · provider-agnostic slightly complex (lean runtime, learning loop, disk-first memory) Built-in skills
3 nanobot 46.5k Ultra-lightweight, self-hosted personal agent framework: the harness is a Python daemon wiring tools, memory, and MCP into chat/webhook front ends (Telegram, Discord, web); minimal footprint alternative to heavier personal-runtime stacks. mcp · memory · local · python mostly simple (lightweight daemon, chat/MCP) Project README
4 CowAgent 46.3k Self-hosted harness (formerly chatgpt-on-wechat) that plans tasks, runs tools/skills, and self-evolves via memory; multi-model, multi-channel (WeChat, Telegram, etc.), one-line install. memory · python slightly complex (multi-channel, self-evolving) Project README
5 Khoj ★ 36.2k Self-hostable "AI second brain": answers over your docs and the web, custom agents, scheduled automations, and multi-client reach (web, Obsidian, Emacs, WhatsApp). A personal-agent harness with retrieval at the core. python complex (server + clients — product suite) Feature tour
6 Eliza ★ 18.9k Open "agentic operating system" (elizaOS): persistent multi-agent runtime with character files, a plugin ecosystem, and social/platform integrations — the harness behind a large share of autonomous social agents. memory · multi-agent · typescript complex (runtime + plugin ecosystem — product suite) Agent quickstart
7 Agent Zero 18.7k Organic, prompt-defined personal agent framework: hierarchical sub-agents, persistent memory, browser and code tools, and self-modifying behavior; runs in Docker with a web UI. memory · multi-agent · browser · sandbox · python slightly complex (prompt-defined, Docker + web UI) Framework tour
8 OpenHarness (HKUDS) 15.2k Open agent harness with a built-in personal agent ("Ohmo") that runs across Feishu, Slack, Telegram, and Discord; core tool-use, skills, memory, multi-agent coordination with auto-compaction for multi-day sessions. memory · multi-agent complex (personal agent + multi-channel — product suite) harness-eval skill
9 AIlice 1.4k Fully autonomous general-purpose agent; one binary, Docker-ready, for when you want "set goal and walk away" without a framework. sandbox · python slightly complex (autonomous, one binary) Task showcase
10 Talon ★ 71 Multi-platform personal agent living in Telegram, Discord, Teams, and the terminal. The harness is a pluggable-backend loop (Claude, Kilo, OpenCode, Codex, OpenAI Agents) with full MCP tool access and persistent background agents (Goals, Heartbeat, Dream); the chat apps are shells. mcp · memory · cli · typescript slightly complex (multi-platform, pluggable backends, MCP) Multi-platform setup

Frameworks

General-purpose agent and LLM application frameworks (the app layer, not harnesses per se).

# Project ⭐ Stars Description Open source Simplicity ↔ capability Examples
1 n8n ★ ✱ 199k Fair-code workflow engine with 400+ nodes and native AI nodes; the self-hosted Zapier that actually does agents and LangChain. workflow · local · typescript ⚠️ Fair-code complex (400+ nodes, workflow engine — product suite) Agent vs chain workflow
2 AutoGPT ★ 186k The original autonomous loop: goal in, agent iterates with tools and memory; Forge is the dev framework, Benchmark the eval harness. memory · evals · python ⚠️ Polyform-SU complex (autonomous loop, tools, memory — product suite) Medium blogger graph
3 langflow ★ 153k Low-code UI to build and deploy LangChain/LangGraph flows; visual DAG editor and one-click run. low-code · python complex (low-code, visual — product suite) Chat with RAG flow
4 Dify ★ 151k One-stop LLM app platform: visual workflows, RAG pipeline, 50+ tools, model management; "ship from prototype to prod" in a single UI. low-code · rag · python ⚠️ Fair-code complex (one-stop platform — product suite) Customer-service bot
5 langchain 143k Chains, tools, retrievers, and agents; the usual entry point for "add tools to an LLM" in Python/JS. python complex (kitchen-sink ecosystem — product suite) Build an agent notebook
6 browser-use 108k Python layer over Playwright: natural-language goals become browser actions—web-agent loop without hand-rolling MCP or a custom driver for every site. mcp · browser · python slightly complex (LLM + browser, Playwright) Grocery shopping agent
7 Flowise ★ 55.1k Drag-and-drop LangChain UI; deploy flows without code. The low-code sibling to Langflow, with a different component and hosting story. low-code · typescript ⚠️ Apache+CLA complex (low-code, drag-drop — product suite) Agentic RAG flow
8 llama-index 51.3k Data-centric: indexing, RAG, and query engines; agent abstractions sit on top of your data pipelines. rag · python complex (RAG + agents — product suite) Research assistant workflow
9 agno 41.5k Python agents with memory, knowledge bases, tools, and structured outputs; continues the PhiData-era product line under the Agno name—production apps, evals, and pipelines. memory · evals · python complex (memory, KB, observability — product suite) Agent with tools
10 langgraph ★ ✱ 38.7k State-machine graphs over LLM steps; checkpointing, human-in-the-loop, and durable execution so workflows survive restarts. workflow · python slightly complex (graphs, checkpointing, durable exec) Customer support agent
11 semantic-kernel 28.4k Microsoft's plugin and planner layer for LLMs; C#, Python, Java; strong on enterprise auth and orchestration. python complex (enterprise, multi-language — product suite) Chat completion agent
12 mastra ✱ 26.8k TypeScript-first; agents, tools, and workflows with a single runtime and minimal boilerplate. typed · typescript ⚠️ Elastic-2.0 slightly complex (TS-first, minimal boilerplate) Durable research agent
13 Haystack 26.1k Open-source orchestration framework for context-engineered LLM apps: modular pipelines and agent workflows with explicit control over retrieval, routing, memory, and generation—closer to LangChain's territory than a coding-agent harness. memory · rag · python complex (modular pipelines, RAG + agents — product suite) Project README
14 letta ★ ✱ 24.1k Python agent runtime with tool use and control flow; lean API; stateful agents with long-horizon memory. memory · python mostly simple (lean API) Loop .af agent file
15 rasa ★ 21.3k Conversational AI stack (NLU, dialogue, actions); long-standing OSS choice for chat and voice bots. voice · python complex (full stack — product suite) Sara conversational demo
16 Google ADK ★ 21k Google's official Agent Development Kit: code-first Python toolkit for building, evaluating, and deploying agents. Optimized for Gemini but model-agnostic; deploys to Cloud Run / Vertex AI; ships a dev UI with eval and a code-execution sandbox. evals · sandbox · python complex (official Google SDK, eval, deploy — product suite) Travel concierge agent
17 botpress ★ 14.8k Visual bot builder and runtime; multi-channel, open-source alternative to commercial bot platforms. low-code · typescript complex (visual builder, multi-channel — product suite) Inter-bot delegation
18 R2R ★ 8k RAG-first: hybrid search, knowledge graphs, multimodal; the framework for "production RAG" when you care more about retrieval than chat UI. vision · rag · workflow · python complex (production RAG — product suite) hello_r2r RAG example
19 agent-squad 7.7k AWS-originated orchestrator (now under 2FastLabs): intent classification, streaming, SupervisorAgent; "agent-as-tools" so one agent delegates to a squad. multi-agent slightly complex (squad orchestration) E-commerce support sim
20 AgentVerse ★ 5.1k Task-solving and simulation envs for multi-LLM agents; deploy many agents in custom environments without building infra from scratch. multi-agent · python complex (simulation envs, multi-agent — product suite) NLP classroom sim
21 youtu-agent 4.6k Tencent Cloud's agent framework: a minimal tool-calling harness designed to perform well with open-source models, positioned as a lighter alternative to heavier orchestration frameworks. mostly simple (minimal loop, open-model focus) Project README
22 Bee Agent Framework 3.3k Python + TypeScript, LF AI–backed; MCP/ACP, workflows, Requirement Agent; the one that pushes "production multi-agent" without LangChain. mcp · multi-agent · python · typescript complex (production multi-agent — product suite) ReAct agent example

This HTML preview is truncated for page performance. The canonical Markdown file contains the complete snapshot.

MARKDOWN METRICS
10618words
35headings
641links
3code blocks
MDRSS ASSESSMENT
Evidence46/100high confidence
Why MDRSS assigned this score
  • Production catalog audit 2026-08-04
  • Taxonomy classified from title, annotation, source and Markdown signals
  • Agent usefulness evaluated from structure, procedures, examples, evidence and retrieval value
Evidence (1)

Discussion 0

Sign in to join the discussion.