AI Coding Agents β€” 8-Platform Comparison

Latest recorded claim verification: 2026-09-09 (4/82 tracked claims). No edition-wide verified-through date is established. Dates and limitations

Verify an agent’s result: bounded bugfix Β· test the tests Β· parallel tasks Β· methodology and runbook

Executive
Deep
β–Έ Filters β–Έ Lenses β–Έ Legend
Platforms:
Lenses:
βœ… Supported⚠️ Partial❌ Nouniqueexperimentalfreedotted = click or hover for details confidence (click)

Choose a next step. Check its basis.

This publication helps you evaluate coding tools and verify their results. Recommendations appear only when current scoped evidence passes every mandatory condition; unresolved inputs remain visibly withheld.

  • Choose a tool β€” define the task and mandatory conditions.
  • Improve a workflow β€” reproduce and check one bounded bugfix.
  • Understand changes β€” inspect recorded changes and their evidence; reader-specific action notes are still being prepared.
Claim review schedules: 30/82 Current 34/82 Due soon 0/82 Overdue 18/82 Withheld Due soon = within 7 days; overdue starts on the due date. Current means within its recorded schedule, not independently proven. Withheld includes stale, disputed or incomplete records.

Report changelog β€” what changed in this living report (not vendor changelogs)

  • 2026-09-10fixStrongDM attribution and OpenCode advisory scope corrected
    StrongDM’s β‰₯US$1,000/day/human-engineer amount is now an attributed author guideline, not measured spend, TCO or productivity; its Cursor YOLO reference is identified as a December 2024 narrative and the current foundation remains unverified. OpenCode GHSA-vxw4-wv6m-9hhh now states the affected range below 1.0.216, patched version 1.0.216, Firefox/user-interaction context and CVSS 8.8 as separate scoped assertions. If evaluating OpenCode, check the selected version: affected versions fail the named criterion, an unknown version remains unconfirmed, and a patched version is not proof of general safety. Publication pilot structure is now 22 of 51 bound targets with 28 atomic assertions; this is not a truth percentage or a full-report audit.
  • 2026-09-03dataR7.1: Plugins and stale capability cells
    Corrected plugin availability and related documented capability gaps. Preserved surface limits and preview labels; unsupported memory and isolation assertions are explicitly unverified. The bounded audit and exact source decisions are retained in project research.
  • 2026-09-03featureFile-backed delivery lifecycle pilot
    Added strict WorkOrder, inspection, plan, approval, checkpoint, verification-link, review, release, and retrospective records for bounded software-delivery work. The first brownfield pilot records restart state and invalidates approval evidence when material inputs change; production release still requires a separate just-in-time owner approval.
Show 16 earlier entries
  • 2026-09-03featurePoint-of-use publication evidence pilot
    Added a fail-closed publication contract for every summary-card bullet, risk item, and Enterprise Adoption case. Twenty-two of 51 targets now show 26 atomic claim bindings with status, verification date, source class, and source link at point of use; the other 29 targets are visibly labelled unbound legacy projections rather than implied as fully sourced. This is a bounded pilot, not completion of report-wide claim binding.
  • 2026-09-03dataSeptember candidate and freshness refresh
    Reviewed 63 source-monitoring candidates and re-verified all 18 active claims due September 7. Updated stable releases and repository counts; Antigravity's model roster; Claude Fable 5.1; Copilot models, approvals and exclusions; Cursor self-hosted workers and Origin semantics; and Kiro Web/CLI/IDE surfaces. Community, prerelease, superseded and out-of-scope signals remain explicitly deferred.
  • 2026-08-24dataEarly-September freshness wave
    Re-verified 25 active claims due in early September against current primary evidence. Updated Codex CLI 0.149.1, OpenCode v1.18.22 and 200,965 stars, Cline v4.1.15 and 66,771 stars, and Kiro GPT-5.6 Terra/Luna credit reductions. Twenty claims remained supported without wording changes; editorial, unrelated ADK, and community-forum signals were deferred.
  • 2026-08-23dataEight-platform primary-source refresh
    Re-audited all eight platforms. Corrected report scope (8, not 9), Codex TOML custom agents, OpenCode CVE-2026-22812 (CVSS 8.8; fixed 1.0.216), current Cline/OpenCode releases, Antigravity and Kiro model rosters, Kiro subagents/advisories, Copilot handoff boundaries, and Claude hook/surface wording. Replaced sixteen broken v6.3 snapshot references; unsupported Antigravity knowledge/security and Cursor CVE details remain stale/withheld. Added one Cline subagent claim and deferred two noisy acquisition signals.
  • 2026-08-21dataHigh-volatility evidence refresh β€” pricing, security and model availability
    Re-researched exactly 12 overdue high-volatility claims against primary sources. Reactivated 10 after substantive rewrites, kept GPT-5.6 stale pending official Sol-price reconciliation, and retired the duplicate Antigravity $100 tier claim. Corrected the HalluSquatting arXiv ID and rate interpretation; replaced Copilot premium requests with AI Credits; reconciled Antigravity tiers; updated current pricing rows, Bugbot billing/metrics, and DuneSlide CVSS wording.
  • 2026-06-19dataKiro added as 9th platform
    Added Kiro (AWS agentic IDE, kiro.dev) as the 9th compared platform across all sections: feature matrix, config, onboarding, execution, skills, subagents (kiro_default), cost (credit-pool pricing), strengths/gaps, summary card, the Init Playbook, and two risk items (workspace-trust CVE cluster β€” high; contested autonomous-agent outage β€” medium). Scorecard + takeaways recompute: Kiro overall 64% (rank 3 of 9). 8 new claim-kiro-* claims with kiro.dev/AWS provenance; two adversarial reviews returned SHIP.
  • 2026-06-12featureInit Playbook + per-cycle changelog
    Added new Project Initialization section with a universal manual init prompt (3-phase scan/draft/curate) and per-platform copy targets for all 8 agents. Added per-cycle changelog block at the bottom of Takeaways. Standalone reference doc (coding-agent-project-init.html) updated to match with corrections from review (OpenCode CVE numbers, Devin Desktop rebrand, Codex /init, etc.).
  • 2026-06-12dataAntigravity 2.0 ingest (Google I/O 2026, 2026-05-19) 29a0773
    Full pipeline cycle for Antigravity 2.0: 8 new candidates, 9 adjudications, 7 new claims (+ 3 bumped), new risk_item for 2.0 release-quality regressions, replaced generic blog.google source with 4 targeted Antigravity sources. Pricing claim marked stale pending tier reconciliation.
  • 2026-05-21featureDecision wizard + freshness + glossary + research-pending 28b24cc
    7 UI/data additions: freshness strip atop Takeaways, cycle delta block above What's New, 15-term glossary, 4-question decision wizard (scored from registry data only), research-pending section with 9 unsourced topics, explicit risk severity + prevalence on all 12 risks, pricing as-of pills from claims.json.
  • 2026-05-04contentFilled 86 missing tooltips a3840d0
    Backfilled execution/config/skills section tooltips to bring tooltip coverage above 95%.
  • 2026-05-04fixTooltip click + tap + hover 85d6767
    Replaced native title attribute with click+tap+hover popover. Native title doesn't fire on click or touch; the new popover does.
  • 2026-05-04dataMay 4 cycle β€” Opus 4.7 + GPT-5.5 wave 610c030
    Refreshed model references and tooltip entity escaping fix.
  • 2026-04-09dataApril 9 fetch cycle 7e0eee1
    6 material platform updates, 4 noise filtered out.
  • 2026-06-29dataWindsurf removed from report (8 platforms now)
    Cognition acquired Windsurf and rebranded to 'Devin Desktop' on 2026-06-02; the original product no longer exists as an independent coding agent. Removed Windsurf across all registry sections (matrix, config, onboarding, execution, skills, cost, strengths, summary_cards, subagents, risks), 3 claims + links, the Init Playbook card, and the cross-platform risk relabeled 'all 9 platforms' β†’ 'all 8 platforms'. Source monitor entry commented out in sources.yaml; historical adjudications/candidates/snapshots preserved.
  • 2026-06-29dataJun 29 fetch cycle β€” Cline 4.0, Cursor 3.9, Opus 4.8, OpenCode fix, enterprise refresh
    Fetched all 8 platform sources. Material: Cline v4.0.0 (Plugins, Customize marketplace, ClinePass, SDK runtime); Cursor 3.8/3.9 (Customize page, Marketplace leaderboard, Automations); Opus 4.8 cross-platform (Claude Code 4.7->4.8, Copilot +4.8, Cursor +4.8 β€” Cline unconfirmed, Antigravity not yet); enterprise refresh (Stripe 1.3k PR/wk, Ramp >50%, Coinbase Forge+Mux, +Cloudflare +Browserbase, 15 in-house agents). Correctness fix: OpenCode 'archived Sep 2025' relabeled active β€” that archive was the old Go repo (now Crush); active OpenCode is anomalyco TS rewrite (v1.17.11). Deferred as noise: Antigravity blog (already-ingested I/O 2026), GPT-5.5 (already tracked), Codex CLI alpha bump, Copilot PR-merge metric.
  • 2026-07-15dataJul 15 fetch cycle β€” GPT-5.6 wave, security (DuneSlide + HalluSquatting), platform version bumps
    Fetched all sources + cross-cutting research. GPT-5.6 Sol/Terra/Luna (Codex/Copilot/Kiro); Anthropic Sonnet 5 + Fable 5; Grok 4.5 in Cursor. Security: Cursor DuneSlide (CVE-2026-50548/50549, patched pre-3.0) + cross-platform HalluSquatting. Version bumps: CC v2.1.209, Cursor 3.11, Cline 4.0.8, Kiro IDE 1.0.138, OpenCode 1.17.20, Antigravity CLI 1.1.2. Usage-based pricing shift (Copilot AI Credits, Cursor repricing). New entrants (Grok Build, JetBrains Junie) + protocol shifts (ACP, MCP 2026-07-28) recorded as research watch-items, not added as platforms.

Decide for your task

Solo developer

One local bounded bugfix with a reproducible verification command.
No forced winner

Preference evidence ranges overlap among possible leaders: Antigravity Β· Claude Code Β· Kiro Β· OpenAI Codex

Task: One local bounded bugfix with a reproducible verification command.

  • Claude Code: eligible; preference range 2–3
  • GitHub Copilot: insufficient evidence; preference range 0–0
  • Cursor: insufficient evidence; preference range 0–0
  • Cline: insufficient evidence; preference range 0–0
  • OpenAI Codex: eligible; preference range 2–3
  • OpenCode: insufficient evidence; preference range 0–0
  • Antigravity: eligible; preference range 2–3
  • Kiro: eligible; preference range 2–3
  • Antigravity Β· Documented local or CLI surface: confirmed pass β€” A dedicated current CLI is documented.
  • Claude Code Β· Documented local or CLI surface: confirmed pass β€” The current claim documents terminal and local-client surfaces.
  • Cline Β· Documented local or CLI surface: unconfirmed β€” Evidence is missing; absence is not treated as incompatibility.
  • Cursor Β· Documented local or CLI surface: unconfirmed β€” Evidence is missing; absence is not treated as incompatibility.
  • GitHub Copilot Β· Documented local or CLI surface: unconfirmed β€” Evidence is missing; absence is not treated as incompatibility.
  • Kiro Β· Documented local or CLI surface: confirmed pass β€” The current claim documents the CLI as a GA surface.
  • OpenAI Codex Β· Documented local or CLI surface: confirmed pass β€” A current stable CLI release is documented.
  • OpenCode Β· Documented local or CLI surface: unconfirmed β€” Evidence is missing; absence is not treated as incompatibility.
  • Antigravity Β· Documented included usage allowance: confirmed pass; weight 2 β€” A $0 Individual tier is documented with plan-dependent limits.
  • Antigravity Β· Open-source client required: unconfirmed; weight 1 β€” Evidence is missing; absence is not treated as incompatibility.
  • Claude Code Β· Documented included usage allowance: confirmed pass; weight 2 β€” Claude Code access is included in named plans, subject to plan limits.
  • Claude Code Β· Open-source client required: unconfirmed; weight 1 β€” Evidence is missing; absence is not treated as incompatibility.
  • Kiro Β· Documented included usage allowance: confirmed pass; weight 2 β€” Named monthly plans include explicit credit allowances.
  • Kiro Β· Open-source client required: unconfirmed; weight 1 β€” Evidence is missing; absence is not treated as incompatibility.
  • OpenAI Codex Β· Documented included usage allowance: confirmed pass; weight 2 β€” Codex access is included in named plans, subject to plan limits.
  • OpenAI Codex Β· Open-source client required: unconfirmed; weight 1 β€” Evidence is missing; absence is not treated as incompatibility.
Evidence and dates

Next step: Resolve the listed preference evidence gaps or compare the possible leaders directly.

Evaluated 2026-09-10 Β· Report-derived guidance for a declared task and scope; not a vendor endorsement or a universal ranking.

Small team

Produce a reviewable pull request that follows team conventions.
No forced winner

Preference evidence ranges overlap among possible leaders: Claude Code Β· Cursor Β· GitHub Copilot

Task: Produce a reviewable pull request that follows team conventions.

  • Claude Code: eligible; preference range 3–3
  • GitHub Copilot: eligible; preference range 1–3
  • Cursor: eligible; preference range 1–3
  • Cline: insufficient evidence; preference range 0–0
  • OpenAI Codex: insufficient evidence; preference range 0–0
  • OpenCode: insufficient evidence; preference range 0–0
  • Antigravity: insufficient evidence; preference range 0–0
  • Kiro: insufficient evidence; preference range 0–0
  • Antigravity Β· Documented change-review workflow: unconfirmed β€” Evidence is missing; absence is not treated as incompatibility.
  • Claude Code Β· Documented change-review workflow: confirmed pass β€” A local code-review command is documented with explicit managed-service boundaries.
  • Cline Β· Documented change-review workflow: unconfirmed β€” Evidence is missing; absence is not treated as incompatibility.
  • Cursor Β· Documented change-review workflow: confirmed pass β€” Security Reviewer checks pull requests in the named Cloud Agent surface.
  • GitHub Copilot Β· Documented change-review workflow: confirmed pass β€” The claim documents code-review and PR-readiness surfaces with restrictions.
  • Kiro Β· Documented change-review workflow: unconfirmed β€” Evidence is missing; absence is not treated as incompatibility.
  • OpenAI Codex Β· Documented change-review workflow: unconfirmed β€” Evidence is missing; absence is not treated as incompatibility.
  • OpenCode Β· Documented change-review workflow: unconfirmed β€” Evidence is missing; absence is not treated as incompatibility.
  • Claude Code Β· Documented local or CLI surface: confirmed pass; weight 2 β€” The current claim documents terminal and local-client surfaces.
  • Claude Code Β· Documented included usage allowance: confirmed pass; weight 1 β€” Claude Code access is included in named plans, subject to plan limits.
  • Cursor Β· Documented local or CLI surface: unconfirmed; weight 2 β€” Evidence is missing; absence is not treated as incompatibility.
  • Cursor Β· Documented included usage allowance: confirmed pass; weight 1 β€” Named plans document included model pools, with on-demand usage separate.
  • GitHub Copilot Β· Documented local or CLI surface: unconfirmed; weight 2 β€” Evidence is missing; absence is not treated as incompatibility.
  • GitHub Copilot Β· Documented included usage allowance: confirmed pass; weight 1 β€” A limited included allowance is documented.
Evidence and dates

Next step: Resolve the listed preference evidence gaps or compare the possible leaders directly.

Evaluated 2026-09-10 Β· Report-derived guidance for a declared task and scope; not a vendor endorsement or a universal ranking.

Enterprise / platform

Evaluate a rollout against a named contractual, residency, retention, certification, or policy requirement.
More scope is required

Name the exact requirement: compliance_scope

Task: Evaluate a rollout against a named contractual, residency, retention, certification, or policy requirement.

Next step: Provide the named missing scope before comparing candidates.

Evaluated 2026-09-10 Β· Report-derived guidance for a declared task and scope; not a vendor endorsement or a universal ranking.

AI / agent researcher

Run isolated repeated research tasks with inspectable restrictions and costs.
Cline

Confirmed match within the declared mandatory scope.

Task: Run isolated repeated research tasks with inspectable restrictions and costs.

  • Claude Code: insufficient evidence; preference range 0–0
  • GitHub Copilot: insufficient evidence; preference range 0–0
  • Cursor: insufficient evidence; preference range 0–0
  • Cline: eligible; preference range 0–1
  • OpenAI Codex: insufficient evidence; preference range 0–0
  • OpenCode: insufficient evidence; preference range 0–0
  • Antigravity: insufficient evidence; preference range 0–0
  • Kiro: insufficient evidence; preference range 0–0
  • Antigravity Β· Isolated parallel research tasks: unconfirmed β€” Evidence is missing; absence is not treated as incompatibility.
  • Antigravity Β· Open-source client required: unconfirmed β€” Evidence is missing; absence is not treated as incompatibility.
  • Claude Code Β· Isolated parallel research tasks: unconfirmed β€” Evidence is missing; absence is not treated as incompatibility.
  • Claude Code Β· Open-source client required: unconfirmed β€” Evidence is missing; absence is not treated as incompatibility.
  • Cline Β· Isolated parallel research tasks: confirmed pass β€” Parallel isolated research is documented with restrictive tool boundaries.
  • Cline Β· Open-source client required: confirmed pass β€” The current claim documents the client license.
  • Cursor Β· Isolated parallel research tasks: unconfirmed β€” Evidence is missing; absence is not treated as incompatibility.
  • Cursor Β· Open-source client required: unconfirmed β€” Evidence is missing; absence is not treated as incompatibility.
  • GitHub Copilot Β· Isolated parallel research tasks: unconfirmed β€” Evidence is missing; absence is not treated as incompatibility.
  • GitHub Copilot Β· Open-source client required: unconfirmed β€” Evidence is missing; absence is not treated as incompatibility.
  • Kiro Β· Open-source client required: unconfirmed β€” Evidence is missing; absence is not treated as incompatibility.
  • OpenAI Codex Β· Isolated parallel research tasks: unconfirmed β€” Evidence is missing; absence is not treated as incompatibility.
  • OpenAI Codex Β· Open-source client required: unconfirmed β€” Evidence is missing; absence is not treated as incompatibility.
  • OpenCode Β· Isolated parallel research tasks: unconfirmed β€” Evidence is missing; absence is not treated as incompatibility.
  • Cline Β· Documented local or CLI surface: unconfirmed; weight 1 β€” Evidence is missing; absence is not treated as incompatibility.
Evidence and dates

Next step: Verify the named surface and plan in one bounded task before adoption.

Evaluated 2026-09-10 Β· Report-derived guidance for a declared task and scope; not a vendor endorsement or a universal ranking.

Decision check β€” hard requirements before preferences

Concrete task profile
Named compliance or contractual requirement
Client license and inference-cost contract
Security requirement and selected OpenCode case
Answer all four scoped questions. Unknown evidence can only produce an insufficient-evidence result.

Inspect the legacy comparison evidence for the underlying rows. Historical coverage remains contextual evidence, not a recommendation.

Compare: legacy coverage details

These historical status counts have not completed per-cell semantic evidence adjudication. They are not tool quality, safety or task recommendations.

Inspect legacy accounting and all underlying rows
Dimension Claude Code Copilot Cursor Kiro Codex Antigravity OpenCode Cline
Features 60% 65% 63% 65% 52% 60% 57% 42%
β–Έ breakdown
Claude Code
12βœ… 5⚠️ 7❌
Copilot
13βœ… 5⚠️ 6❌
Cursor
12βœ… 5⚠️ 6❌
Kiro
12βœ… 7⚠️ 5❌
Codex
9βœ… 7⚠️ 8❌
Antigravity
14βœ… 1⚠️ 9❌
OpenCode
9βœ… 8⚠️ 6❌
Cline
8βœ… 4⚠️ 12❌
Only Claude Code: Peer-to-peer comms
Only Antigravity: Browser subagent
Only Antigravity: Commentable artifacts
Customization 83% 89% 75% 83% 56% 78% 0% 50%
β–Έ breakdown
Claude Code
6βœ… 3⚠️ 0❌
Copilot
7βœ… 2⚠️ 0❌
Cursor
5βœ… 2⚠️ 1❌
Kiro
6βœ… 3⚠️ 0❌
Codex
2βœ… 6⚠️ 1❌
Antigravity
6βœ… 2⚠️ 1❌
OpenCode
0βœ… 0⚠️ 1❌
Cline
2βœ… 5⚠️ 2❌
Only Copilot: Agent handoffs
Only Copilot: Knowledge Base
Skills 100% 50% 50% 50% 75% 50% 100% 50%
β–Έ breakdown
Claude Code
2βœ… 0⚠️ 0❌
Copilot
1βœ… 0⚠️ 1❌
Cursor
1βœ… 0⚠️ 1❌
Kiro
1βœ… 0⚠️ 1❌
Codex
1βœ… 1⚠️ 0❌
Antigravity
1βœ… 0⚠️ 1❌
OpenCode
2βœ… 0⚠️ 0❌
Cline
1βœ… 0⚠️ 1❌
Onboarding 40% 75% 75% 50% 60% 50% 60% 70%
β–Έ breakdown
Claude Code
2βœ… 0⚠️ 3❌
Copilot
3βœ… 0⚠️ 1❌
Cursor
3βœ… 0⚠️ 1❌
Kiro
3βœ… 1⚠️ 3❌
Codex
3βœ… 0⚠️ 2❌
Antigravity
2βœ… 1⚠️ 2❌
OpenCode
3βœ… 0⚠️ 2❌
Cline
3βœ… 1⚠️ 1❌
Only Cursor: Setup script
No platform fully supports: Init command (partial only)
Execution 77% 79% 71% 73% 69% 65% 58% 54%
β–Έ breakdown
Claude Code
9βœ… 2⚠️ 2❌
Copilot
9βœ… 1⚠️ 2❌
Cursor
8βœ… 1⚠️ 3❌
Kiro
7βœ… 5⚠️ 1❌
Codex
8βœ… 2⚠️ 3❌
Antigravity
7βœ… 3⚠️ 3❌
OpenCode
5βœ… 4⚠️ 3❌
Cline
5βœ… 3⚠️ 4❌
Only Cursor: Model comparison
Only Antigravity: Commentable artifacts
Overall (provisional) 72% 72% 67% 64% 62% 61% 55% 53%

Overall = simple mean of the 5 dimension percentages. Each dimension percentage = (βœ… count + 0.5 Γ— ⚠️ count) / row count, scored only over rows we track. Several enterprise-decisive dimensions (performance benchmarks, telemetry/privacy, compliance, model coverage, IDE/OS support, failure modes) are not yet scored β€” read the Overall row as "highest coverage in this matrix today", not as a quality verdict.

Enterprise Adoption

Real-world deployment examples. These are market signals, not normalized feature scores.

Stripe β€” Minions

Foundation: Goose (Block fork)
Scale: 1,300+ PRs/week
Invocation: Slack, CLI, Web
Human review: required
Zero human-written code. Devboxes spin up in 10s. Blueprints deterministic orchestration.
Bound assertion: β€œ1,300+ PRs/week”Withheldclaim-stripe-minions-scale Β· Ry Walker: In-House Coding Agents research (secondary Β· manual_research) Β· recorded verification 2026-06-29 Β· review due 2026-07-13

Spotify β€” Honk

Foundation: Claude Code
Scale: 50+ features shipped
Invocation: Slack
Human review: not disclosed
Best developers haven't written code since December 2025. Stock +14.7% on earnings.
Bound assertion: β€œ50+ features shipped”Withheldclaim-spotify-claude-code-adoption Β· Ry Walker: In-House Coding Agents research (secondary Β· manual_research) Β· recorded verification 2026-05-04 Β· review due 2026-06-03

Ramp β€” Inspect

Foundation: OpenCode
Scale: >50% of merged PRs
Invocation: Slack, Web, Chrome ext, Voice, Mobile
Human review: required
Now highest reported adoption (>50%), surpassing Abnormal's 13%. Organic uptake; warm sandbox preloading.
Bound assertion: β€œ>50% of merged PRs”Withheldclaim-ramp-opencode-adoption Β· Ry Walker: In-House Coding Agents research (secondary Β· manual_research) Β· recorded verification 2026-06-29 Β· review due 2026-07-13

Coinbase β€” Forge (formerly Claudebot/Cloudbot)

Foundation: Multi-model (model-agnostic)
Scale: 5% of PRs; Mux: 3.5x throughput, 600+ engineers, 5,068 PRs/461 repos
Invocation: Slack, Linear
Human review: context-dependent
150h β†’ 15h PR cycle time. New Mux fleet-orchestration layer runs parallel agent fleets. Built by VP Chintan Turakhia; deep Skills/MCP integration (Datadog, Sentry, Amplitude, Snowflake).
Bound assertion: β€œ3.5x throughput”Withheldclaim-coinbase-cycle-time Β· Ry Walker: In-House Coding Agents research (secondary Β· manual_research) Β· recorded verification 2026-06-29 Β· review due 2026-07-13

Shopify β€” Roast

Foundation: Claude Code CLI
Scale: not disclosed
Invocation: CLI
Human review: required
Open-source Ruby gem (roast-ai). DSL for structured AI workflows.
Unbound legacy projection Β· Legacy secondary-source case has no dedicated atomic claim in the current registry.

Uber β€” Validator + Autocover + Picasso

Foundation: LangGraph
Scale: 21,000 dev hours saved
Invocation: IDE, Workflow
Human review: not disclosed
5,000 developers, hundreds of millions LOC. 10% test coverage increase.
Bound assertion: β€œ21,000 dev hours saved”Withheldclaim-uber-dev-hours-saved Β· Ry Walker: In-House Coding Agents research (secondary Β· manual_research) Β· recorded verification 2026-03-18 Β· review due 2026-05-18

Abnormal AI β€” Internal agents

Foundation: Custom
Scale: 13% of PRs
Invocation: not disclosed
Human review: not disclosed
13% of PRs from background agents (Feb 2026); since surpassed by Ramp (>50%). Migrating to Modal.
Unbound legacy projection Β· Legacy secondary-source case has no dedicated atomic claim in the current registry.

StrongDM β€” Software Factory

Foundation: Cursor YOLO in a December 2024 historical narrative; current foundation unverified
Scale: Author guideline: β‰₯US$1,000/day/human engineer in tokens
Invocation: Spec-driven
Human review: Author-described: no human code review
Justin McCarthy self-report; no current platform foundation, safety, measured spend, TCO, productivity, or external validation is established.
Bound assertion: β€œAuthor guideline: β‰₯US$1,000/day/human engineer in tokens”Currentclaim-strongdm-token-guideline Β· Justin McCarthy, Software Factory: Finally, in practical form (official Β· official_updates) Β· recorded verification 2026-09-09 Β· review due 2026-10-09

Cloudflare β€” AI Gateway + multi-agent Code Reviewer (harness, not agent)

Foundation: In-house harness (OpenCode/Windsurf assistant layer; Workers + Durable Objects)
Scale: 3,683 users (93% of R&D); 241B tokens/30d; $1.19/code review
Invocation: IDE assistant + GitLab CI
Human review: augmented (multi-agent reviewer)
"Build the harness, not the agent." MCP Server Portal (13 servers, 182+ tools); auto-generated AGENTS.md across ~3,900 repos; publishes unit economics ($1.19/review).
Unbound legacy projection Β· Legacy secondary-source case is only included in an aggregate stale claim, not an atomic case claim.

Browserbase β€” Internal maintenance agents

Foundation: Custom (internal)
Scale: not disclosed (maintenance-focused)
Invocation: Automated (CI / maintenance)
Human review: not disclosed
Agents-as-maintainers: auto-clean code/config as features ramp down β€” breaks the product-bloat -> LLM-degradation loop (Kyle Jeong).
Unbound legacy projection Β· Legacy secondary-source case is only included in an aggregate stale claim, not an atomic case claim.
Source: Ry Walker Research Β· 2026-06-11

Claude Code

  • Auto: 3 built-in subagents auto-delegate; Chrome browser (beta); cloud VMs; Claude Code Security (AI vuln scanner)
    Unbound legacy projection Β· Compound legacy capability summary has not yet been atomized into claim-backed assertions.
  • Config: Custom agents via .claude/agents/*.md; Agent Teams; plugins; Auto Memory
    Unbound legacy projection Β· Compound legacy configuration summary has not yet been atomized into claim-backed assertions.
  • New: Opus 4.8 GA (May 28) + xhigh effort tier (v2.1.111); native binary distribution (v2.1.113); expanded lifecycle hooks including PreCompact, MCP-tool handlers, and PostToolUse output replacement; /ultrareview cloud multi-agent review; PowerShell tool (Windows preview); Vim visual modes; 1-hour prompt cache TTL; per-model cost breakdown; screen-reader mode (--ax-screen-reader) + Chrome extension; Claude Code 2.1.257 adds Fable 5.1 (1M context) and 2.1.259 adds managed MCP plus fail-closed managed settings; current v2.1.259
    Bound assertion: β€œClaude Code 2.1.257 adds Fable 5.1 (1M context)”Due soonclaim-anthropic-sonnet5-fable5 Β· Claude Code 2.1.257 changelog (official Β· official_updates) Β· recorded verification 2026-09-03 Β· review due 2026-09-17
    Bound assertion: β€œOpus 4.8 GA (May 28)”Due soonclaim-opus-48-platform-availability Β· Claude Code model configuration (official Β· official_docs) Β· recorded verification 2026-09-03 Β· review due 2026-09-17

GitHub Copilot

  • Auto: Agent Mode (local); Coding Agent (50% faster, GH Actions); Cloud Agent (research+plan+code); Copilot SDK (public preview); GPT-5.5 GA (Apr 24) + Opus 4.8 GA (May 28) + GPT-5.6 Sol/Terra/Luna (Jul 9); auto model selection in CLI
    Unbound legacy projection Β· Compound legacy capability summary has not yet been atomized into claim-backed assertions.
  • Config: Custom agents via .github/agents/*.agent.md; hooks; skills; handoffs; BYO LLM key (VS Code); C++ in CLI (preview); US+EU data residency + FedRAMP-authorized models; org runner controls + firewall
    Unbound legacy projection Β· Compound legacy configuration summary has not yet been atomized into claim-backed assertions.
  • New: Sep 2 content exclusions GA in app/CLI; Sep 1 Fable 5.1 GA and opt-in code-review approvals; enterprise/team managed default models; /security-review public preview; usage-based AI Credits pricing (Jun 1); Opus 4.8 GA (May 28); Opus 4.7 GA (Apr 16) + GPT-5.5 GA (Apr 24); GPT-5.2/5.2-Codex deprecation (May 1); Opus 4.6 Fast retired (Apr 10); code review consumes Actions minutes from June 1 2026; paused new Copilot Business self-serve signups (Apr 22)
    Bound assertion: β€œSep 2 content exclusions GA in app/CLI”Due soonclaim-copilot-security-review Β· Copilot app and CLI content exclusions (official Β· official_updates) Β· recorded verification 2026-09-03 Β· review due 2026-09-17

Cursor

  • Auto: Explore/Bash/Parallel subagents; /multitask async subagents (v3.2); BG Agents (8 cloud, self-hosted); Long-Running Agents; Bugbot PR review; Multi-Agent Judging; Security Review beta (always-on PR vuln + prompt-injection scanner)
    Unbound legacy projection Β· Compound legacy capability summary has not yet been atomized into claim-backed assertions.
  • Config: .cursor/agents/*.md + UI modes; .cursor/skills/; Hooks (12 events); Plugin Marketplace + Team Marketplace controls; /worktree + /best-of-n + /multitask; multi-root workspaces (single agent, multiple repos); Cursor SDK (TypeScript public beta); Composer 2 model; Canvases (durable visual artifacts); Customize page (v3.9); supports Claude Opus 4.8 + Grok 4.5 (all plans, Jul 8)
    Unbound legacy projection Β· Compound legacy configuration summary has not yet been atomized into claim-backed assertions.
  • New: Sep 2 β€” self-hosted Linux/macOS workers and dynamic team pools. Aug 19 β€” Cloud Agents subscriptions, isolated subagent VMs, /goal, custom modes, and steering. Aug 17 β€” Origin early beta; Origin-hosted repos use Origin as source of truth while synced GitHub repos keep GitHub authoritative. Historical: v3.11 side chats and conversation search.
    Bound assertion: β€œself-hosted Linux/macOS workers”Due soonclaim-cursor-customize-automations Β· Cursor changelog (official Β· official_updates) Β· recorded verification 2026-09-03 Β· review due 2026-09-17
    Bound assertion: β€œv3.11 side chats and conversation search”Currentclaim-cursor-v311-sidechats Β· Cursor changelog (official Β· official_updates) Β· recorded verification 2026-08-23 Β· review due 2026-11-23

Cline

  • Auto: read-only research subagents run in parallel with isolated context and separate cost tracking; they cannot edit, browse, use MCP, or nest. Plan/Act modes and checkpoints remain available.
    Unbound legacy projection Β· Compound legacy capability summary has not yet been atomized into claim-backed assertions.
  • Config: 100+ models through BYOK or Cline usage-billing; SKILL.md; hooks; VS Code, JetBrains, CLI, ACP, and Kanban. Plugins add tools/hooks/commands on SDK, CLI, and Kanban only (not current VS Code/JetBrains extensions). Optional ClinePass $9.99/month; current stable v4.1.17.
    Bound assertion: β€œPlugins add tools/hooks/commands on SDK, CLI, and Kanban only”Due soonclaim-cline-v4-plugins-marketplace Β· Cline plugins (official Β· official_docs) Β· recorded verification 2026-09-03 Β· review due 2026-09-17
  • New: v4.0.0 (Jun 2026) β€” SDK-backed VS Code runtime; Cline Plugins (custom tools/workflows/skills/MCP); Customize marketplace (install/manage Skills, MCP, Plugins); ClinePass managed provider; reasoning-effort incl. xhigh for DeepSeek; +models GLM 5.2, Kimi K2.7, Qwen 3.7
    Unbound legacy projection Β· Compound legacy release summary has not yet been atomized into claim-backed assertions.

OpenAI Codex

  • Auto: Cloud containers; local worktrees; App Automations (unprompted work)
    Unbound legacy projection Β· Compound legacy capability summary has not yet been atomized into claim-backed assertions.
  • Config: AGENTS.md + config.toml; built-in default/worker/explorer agents; project or personal custom agents in .codex/agents/*.toml; SKILL.md; MCP; open-source Rust CLI; current stable v0.153.0.
    Bound assertion: β€œcurrent stable v0.153.0”Currentclaim-codex-cli-alpha-release Β· Codex CLI 0.153.0 (official Β· official_updates) Β· recorded verification 2026-09-03 Β· review due 2026-12-02
  • New: CLI 0.153.0 (Sep 3); interactive agents dashboard, /agent thread inspection, /cd and /pwd navigation, and codex queue alongside current subagent workflows.
    Unbound legacy projection Β· Release features beyond the bound stable-version assertion are not yet atomized.

OpenCode free MIT free active v1.18.27

  • Auto: Build + Plan agents; General & Explore subagents (parallel); GitHub Agent (Actions-based); hidden context agents
    Unbound legacy projection Β· Compound legacy capability summary has not yet been atomized into claim-backed assertions.
  • Config: 6 extensibility layers: AGENTS.md (CLAUDE.md compat), MCP (OAuth 2.0), npm plugins, a broad documented plugin-event surface, SKILL.md, custom tools; Ollama local models; 6 config precedence levels
    Unbound legacy projection Β· Compound legacy configuration summary has not yet been atomized into claim-backed assertions.
  • New: v1.18.27 (Sep 2); canonical anomalyco/opencode repository remains actively maintained. The archived opencode-ai/opencode repository is a separate historical Go codebase.
    Bound assertion: β€œv1.18.27 (Sep 2)”Currentclaim-opencode-active-status Β· OpenCode v1.18.27 (official Β· official_updates) Β· recorded verification 2026-09-03 Β· review due 2026-10-03

Antigravity free

  • Auto: Browser subagent; dynamic parallel subagents; scheduled tasks; semantic skill activation. Knowledge Subagent specifics are withheld pending a current official locator.
    Unbound legacy projection Β· Compound legacy capability summary includes an explicit withheld sub-claim and requires atomization.
  • Config: GEMINI.md + AGENTS.md rules β†’ Workflows β†’ Skills; Gemini 3.8/3.7/3.6 Flash, Gemini 3.1 Pro, Claude Sonnet/Opus 4.6, GPT-OSS-120b; plan-dependent availability; Terminal Policy controls.
    Unbound legacy projection Β· Compound legacy model and configuration summary has not yet been atomized.
  • New: 2.0 (2026-05-19, Google I/O): separate desktop app without an embedded IDE alongside the Antigravity IDE; new Antigravity CLI and SDK; Antigravity in Gemini Enterprise Agent Platform; project-based context, worktrees, and live voice transcription; current CLI release 1.1.25 (Sep 3, 2026); ADK Go 2.0 multi-agent framework (blog)
    Bound assertion: β€œ2.0 (2026-05-19, Google I/O)”Due soonclaim-antigravity-2-0 Β· Google I/O 2026 developer highlights (official Β· official_updates) Β· recorded verification 2026-09-03 Β· review due 2026-09-17
    Bound assertion: β€œcurrent CLI release 1.1.25”Due soonclaim-antigravity-cli Β· Antigravity CLI 1.1.25 (official Β· official_updates) Β· recorded verification 2026-09-03 Β· review due 2026-09-17
  • Pricing: GA Individual $0; Google AI Pro $20; Ultra $100 with 5x Pro capacity; top Ultra $200 with 20x Pro capacity.
    Unbound legacy projection Β· Pricing belongs to the existing cost claim but is not yet atomized on this duplicate surface.

Kiro

  • Auto: isolated parallel subagents, dependency DAGs, bounded review loops, Autopilot/Supervised modes, agent hooks, and Auto model selection.
    Unbound legacy projection Β· Compound legacy capability summary has not yet been atomized into claim-backed assertions.
  • Config: Custom agents (.kiro/agents/); steering files; SKILL.md skills; MCP; Powers extensions; specs (Requirementsβ†’Designβ†’Tasks)
    Unbound legacy projection Β· Compound legacy configuration summary has not yet been atomized into claim-backed assertions.
  • New: Kiro Web GA (Sep 1); GPT-5.6 Sol/Terra/Luna and Claude Opus 5; Sol 2.4x, Terra 1.0x, Luna 0.1x; CLI 2.21 adds session/config dashboards and cloud sync; IDE 1.0.437 adds cloud configuration and synced Powers. Subagents now support two built-ins, custom IDE/CLI agents, planned DAGs, and bounded review loops.
    Bound assertion: β€œKiro Web GA (Sep 1)”Currentclaim-kiro-ga-release Β· Kiro Web GA (official Β· official_updates) Β· recorded verification 2026-09-03 Β· review due 2026-10-03
    Bound assertion: β€œCLI 2.21 adds session/config dashboards and cloud sync”Due soonclaim-kiro-gpt56-ide138 Β· Kiro CLI 2.21 (official Β· official_updates) Β· recorded verification 2026-09-03 Β· review due 2026-09-17

Quick Feature Matrix

Feature Claude Code Copilot Cursor Cline Codex OpenCode Antigravity Kiro
Custom agent defs βœ… βœ… βœ… ❌ βœ… .codex/agents/*.toml βœ… .opencode/agents/ βœ… Markdown subagents βœ… JSON in .kiro/agents/
Cloud execution ⚠ Web VMs βœ… GH Actions βœ… ❌ βœ… ⚠ Headless + GH Actions ❌ βœ… Kiro Web
Git worktree isolation βœ… ⚠ Branch βœ… ⚠ Checkpoints βœ… Isolation unverified βœ… Optional branch mode ⚠ Cloud isolated envs
Peer-to-peer comms βœ… ❌ ❌ ❌ ❌ ❌ ❌ ❌
Parallel agents βœ… βœ… Coding Agent βœ… βœ… βœ… βœ… Native βœ… βœ… ≀4 subagents
Built-in code review βœ… /code-review βœ… @copilot ⚠ Bugbot ❌ βœ… ⚠ via GitHub Agent ❌ βœ… Supervised mode
Browser subagent ⚠ Chrome ❌ ⚠ ⚠ Built-in ❌ ⚠ webfetch only βœ… ❌
Commentable artifacts ❌ ❌ ❌ ❌ ❌ ❌ βœ… ⚠ Inline hunk chat
Knowledge Base ⚠ Auto Memory βœ… Agentic Memory Unverified automatic memory ⚠ Memory Bank ⚠ Opt-in local memories ⚠ Session-scoped ⚠ withheld ⚠ Steering + knowledge tool
Open source ❌ ⚠ Chat OSS ❌ βœ… Apache-2.0 βœ… CLI βœ… MIT Β· active ❌ ⚠ Code OSS base
Hooks βœ… βœ… βœ… βœ… ⚠ Experimental βœ… plugin events βœ… hooks.json βœ… Agent hooks
Task dependencies βœ… ❌ ❌ ❌ ❌ ⚠ Todo tool ❌ βœ… Dependency graph
Session resumption βœ… βœ… βœ… βœ… βœ… βœ… SQLite βœ… ⚠ History + switching
Agent Skills βœ… βœ… βœ… βœ… βœ… βœ… SKILL.md βœ… βœ… SKILL.md (open std)
Workflows ⚠ via Skills ⚠ Prompts+Skills ⚠ Commands βœ… ⚠ via Skills ⚠ via Skills βœ… βœ… Specs + automations
Multi-vendor models ❌ βœ… Auto+BYOK βœ… βœ… ❌ βœ… broad catalog βœ… plan-dependent βœ… OpenAI + Claude + open
Native IDE ❌ βœ… 6+ IDEs βœ… ⚠ VS Code + JetBrains + CLI ❌ βœ… 6 surfaces βœ… βœ… Code OSS-based
Arena / model A/B ❌ ❌ βœ… Judging ❌ ❌ ❌ ❌ ❌
Agent Manager UI ❌ βœ… Agents Panel βœ… Agents Window ❌ βœ… App ⚠ TUI + web βœ… ⚠ CLI only
Agent-to-agent handoffs ⚠ Teams βœ… guided handoffs ❌ ❌ ❌ ❌ ❌ ⚠ Hierarchical
Voice input βœ… /voice ❌ ❌ ❌ ⚠ Web only ❌ ❌ ❌
Remote / mobile access βœ… Remote Control ⚠ GH Mobile ⚠ Web / iOS beta ❌ ⚠ Web UI ⚠ HTTP server βœ… Remote Control βœ… Web + iOS
AI security scanning βœ… Security ⚠ CodeQL + Autofix βœ… Security Agents ❌ ⚠ Security ❌ ❌ ❌
Free tier ❌ βœ… Free ⚠ βœ… OSS βœ… Free βœ… MIT + free models βœ… Full βœ… $0 Β· 50 credits/mo

Built-in Agents & Subagents

Claude Code 5 agents
AgentCapabilityPurposeToolsModelAccess
Explore Search Search files & codebase Glob, Grep, Read, Bash (read-only) Haiku πŸ”’ Read 3 thoroughness levels: quick Β· medium Β· very thorough
Plan Plan Plan by researching codebase Read, Glob, Grep, Bash Sonnet πŸ”’ Read
General-purpose Execute Execute complex multi-step operations All tools Sonnet ✏ Read/Write
Team Lead experimental Coordinate Coordinate teammates via task list All + SendMessage, TaskList Opus ✏ R/W 2–5 teammates recommended; tested up to 16 experimental
Teammate (Γ—N) experimental Execute Coordinate Execute tasks with peer messaging All + SendMessage (peer-to-peer) Sonnet ✏ R/W Agent Teams experimental
Copilot 5 agents
AgentCapabilityPurposeToolsModelAccess
Plan Mode Plan Plan & approve blueprint before coding search, fetch (read-only) Auto πŸ”’ Plan output Review & approve before agent starts coding
Agent Mode Execute Execute autonomous edits in IDE All IDE tools, terminal, file R/W Claude Sonnet 4.5 ✏ R/W Multi-file editing, terminal, auto-fix. 6+ IDEs
Coding Agent Execute Cloud Execute autonomous cloud dev β†’ PR Full R/W in GH Actions sandbox Claude Sonnet 4.5 ✏ R/W Issueβ†’PR autonomous. Security scanning built-in. Consumes GH Actions minutes + AI Credits
Custom Agent (Γ—N) Execute Execute specialized tasks per profile Configurable via tools property Configurable ✏ R/W .github/agents/*.agent.md; handoffs between agents (IDE only; not yet on GitHub.com); org/enterprise level
Agents Panel Coordinate Coordinate all agent sessions View, steer, resume sessions N/A ⚠ Monitor Agent HQ: mission control for Copilot + 3rd-party agents
Cursor 7 agents
AgentCapabilityPurposeToolsModelAccess
Explore Search Search codebase semantically Semantic search, grep, file read Fast πŸ”’ Read Parallel searches
Bash Execute Execute shell commands Terminal Inherit ⚠ Execute only
Custom Subagent (Γ—N) Execute Execute user-defined parallel tasks Configurable per agent Configurable ✏ R/W .cursor/agents/*.md (v2.4+) + UI modes ; async, nested (v2.5)
Background Agent Execute Cloud Execute autonomous dev in cloud Full R/W in remote Ubuntu VM (AWS) Max Mode ✏ R/W Up to 8 parallel; usage billed; Slack/Linear notify
Long-Running Agent Execute Cloud Autonomous multi-hour/day execution Full R/W, plans + multi-agent verification Max Mode ✏ R/W Ultra/Teams/Enterprise. cursor.com/agents. Research preview
Browser Browse Browse via MCP integration DOM, screenshots, console, network Inherit πŸ”’ Read
Bugbot Execute Review PRs on GitHub GitHub PR analysis, custom rules Dedicated ⚠ Review: new/renewed subscriptions usage billed, average $1-$1.50/run; legacy until renewal. Vendor reports 80% of identified bugs resolved by merge. .cursor/BUGBOT.md rules
Cline 1 agent
AgentCapabilityPurposeToolsModelAccess
Research Subagent (Γ—N) Search Search across codebase in parallel (native use_subagents tool, v3.58.0+) read_file, list_files, search_files, list_code_definition_names, execute_command (ro), use_skill Inherit πŸ”’ Read No writes, no browser, no MCP, no nesting
Codex 4 agents
AgentCapabilityPurposeToolsModelAccess
Cloud Task Execute Cloud Execute parallel tasks in containers Full R/W in sandbox gpt-5.3-codex ✏ R/W (sandboxed) Up to 4 best-of-N; no inter-task comms
Local Worktree Execute Execute independent local tasks Full R/W in worktree Inherit ✏ R/W
MCP Server Coordinate Coordinate by exposing Codex as tool codex + codex-reply via MCP Configurable ⚠ Proxy Multi-agent orchestration via Agents SDK
Automations Execute Config Unprompted scheduled work (issue triage, alerts, CI/CD) Background worktrees or project dir Inherit ✏ R/W Results in Triage inbox
OpenCode 6 agents
AgentCapabilityPurposeToolsModelAccess
Build Plan Execute Full development: plan, code, test, debug All 15 tools: bash, edit, write, read, grep, glob, list, patch, lsp, skill, todo, webfetch, websearch, question Configurable ✏ R/W Primary agent. AGENTS.md instructions. Thinking budget toggle (Ctrl+T)
Plan Plan Read-only analysis & planning Read, Grep, Glob, List, LSP Configurable πŸ”’ Read No writes. Analysis & architecture planning
General (subagent) Execute Parallel multi-step operations All tools (inherited from parent) Configurable ✏ R/W Parallel execution via multiple tool-use blocks
Explore (subagent) Search Fast read-only codebase search Read, Grep, Glob, List Fast model πŸ”’ Read Lightweight, rapid exploration
GitHub Agent Execute Cloud Respond to /opencode on issues & PRs Full R/W in GitHub Actions Configurable ✏ R/W opencode github install β†’ GitHub App. Creates branches, reviews code
Custom Agent (Γ—N) Execute User-defined specialized agents Configurable per agent Configurable ✏ R/W .opencode/agents/*.md or opencode.json; per-agent MCP, tools, model, yolo
Antigravity 2 agents
AgentCapabilityPurposeToolsModelAccess
Main Coding Agent Plan Execute Plan, execute codegen, coordinate All IDE tools, terminal, file R/W Gemini 3.1 Pro ✏ R/W Planning Mode vs Fast Mode; Terminal Policy: Off · Auto · Turbo
Browser Subagent Browse Browse, automate, test, record video Navigate, click, scroll, type, screenshot, JS exec, video Gemini 2.5 CU ⚠ Browser only Isolated Chrome profile; WebP video recording; requires browser extension
Kiro 1 agent
AgentCapabilityPurposeToolsModelAccess
Internal Subagents (Γ—2) Search Execute Context gathering and general-purpose delegated work Shared workspace; isolated context; configured tools and permissions Auto / selected Parallel isolated execution Β· planned DAGs Β· bounded review loops

Custom Agent Configuration

Feature Claude Code Copilot Cursor Cline Codex OpenCode Antigravity Kiro
Custom agents βœ… .claude/agents/*.md βœ… .github/agents/*.agent.md βœ… .cursor/agents/*.md ❌ βœ… .codex/agents/*.toml opencode.json + .opencode/agents/*.md βœ… .agents/agents/*.md βœ… .kiro/agents/ (JSON/MD)
Model selection haiku Β· sonnet Β· opus Β· inherit Auto (default) / model: per agent fast Β· inherit Β· model ID Per-mode Per-profile / --model 75+ providers; Ctrl+T thinking toggle Per-conversation Gemini + Claude + GPT Auto Β· Claude Β· open-weight
Tool restriction βœ… Allow/deny list + MCP βœ… tools: array in YAML βœ… Same + readonly ⚠ ⚠ Sandbox modes Per-agent + per-skill YAML ⚠ βœ… Allow-list + @server
Permission modes βœ… 5 modes ⚠ PR review gated ⚠ Yolo ⚠ --yes ⚠ 3 modes 3 modes (ask / YOLO / per-agent) βœ… 3 levels βœ… Autopilot/Supervised
Agent handoffs ⚠ Teams messaging βœ… handoffs property ❌ ❌ ❌ ❌ No handoffs ❌ ⚠ Parallel; no peer
Marketplace βœ… Plugin marketplace βœ… MCP Registry + awesome-copilot βœ… Plugin Marketplace ⚠ MCP Marketplace ⚠ Plugin directory / CLI npm plugins (@opencode-ai/plugin) βœ… MCP Store βœ… Open VSX + Powers
Alternative Customization
Skills (SKILL.md) βœ… .claude/skills/ context:fork β†’ subagent βœ… .github/skills/ + personal + .claude/ cross-compat βœ… + Commands + Hooks βœ… .cline/skills/ βœ… .agents/skills/ .opencode/skills/ + ~/.config/opencode/skills/ βœ… .agent/skills/ βœ… .kiro/skills/
Rules CLAUDE.md .github/copilot-instructions.md .cursor/rules/ .clinerules/ + AGENTS.md fallback AGENTS.md + config.toml AGENTS.md (6-level precedence) GEMINI.md + .agent/rules/*.md Steering (.kiro/steering/)
Workflows ⚠ via Skills ⚠ Prompts+Skills ⚠ Commands βœ… ⚠ via Skills Custom agents + Skills + hooks βœ… Chainable ⚠ via Specs/Powers
Knowledge Base ⚠ Auto memory βœ… Agentic Memory Unverified automatic memory ⚠ Memory Bank ⚠ Opt-in local memories AGENTS.md (6-level precedence) ⚠ withheld ⚠ Steering + resources
Plugins βœ… /plugin install βœ… plugin.json packages βœ… Plugin bundles ⚠ SDK / CLI / Kanban ⚠ Desktop + CLI npm + hooks + MCP + ACP βœ… plugin.json bundles βœ… Open VSX

Agent Skills Open Standard

Feature Claude Code Copilot Cursor Cline Codex OpenCode Antigravity Kiro
SKILL.md support βœ… Native Creator of standard βœ… Native VS Code, CLI, Coding Agent βœ… Native v2.4+ (Jan 2026) βœ… Native Always-on since v3.57 βœ… Native βœ… YAML frontmatter βœ… Native βœ… Required SKILL.md
Skill location .claude/skills/ Β· ~/ .github/skills/ Β· ~/.copilot/skills/ + .claude/skills/ cross-compat .cursor/skills/ Β· ~/ Β· built-in .cline/skills/ Β· ~/.agents/skills/ + .clinerules/ + .claude/ paths .agents/skills/ Β· ~/ .opencode/skills/ + ~/.config/opencode/skills/ On-demand via built-in skill tool .agent/skills/ .kiro/skills/ Β· ~/
Activation Auto + /skill Auto + /slash command Auto + explicit Auto + use_skill Auto + $skill Pattern + on-demand (glob matching) Semantic auto Auto + /skill
Platform extensions context:fork, $ARGUMENTS Prompts + Instructions + Hooks + Plugins Commands + Hooks Cross-agent paths openai.yaml sidecar npm plugins + MCP + hooks + ACP scripts/ + refs/ Open standard Β· references/
Spawn agents? βœ… context:fork ❌ ❌ ❌ ⚠ βœ… Native multi-agent ❌ ❌
Precedence Project first Repo first Project first Global first CWD β†’ root 6-level hierarchy Workspace first Workspace first

Project Onboarding & Agent Memory

Feature Claude Code Copilot Cursor Cline Codex OpenCode Antigravity Kiro
Instruction Files
Primary file CLAUDE.md .github/copilot-instructions.md .cursor/rules/*.mdc .clinerules AGENTS.md AGENTS.md / CLAUDE.md GEMINI.md Steering files (no single file)
Format Markdown + YAML (rules) Markdown + YAML (.instructions) YAML frontmatter (4 modes) Markdown + YAML (paths) Markdown + Starlark Markdown + JSON Markdown + YAML Markdown (agents JSON)
Hierarchy depth 4 levels + rules + @import 3 levels (accumulate) 3 levels + legacy 2 levels + cross-read N levels (walk-down) 7 sources 3 levels + rules dir 2 levels (workspace > global)
Memory & Persistence
Auto-memory βœ… MEMORY.md (local, per-project) βœ… Server-side (28-day expiry) Unverified automatic memory ⚠ Memory Bank (prompt-driven) βœ… Opt-in local memories ❌ SQLite sessions only ⚠ /memory add (manual) βœ… Foundation always loaded
Ignore, Init & Setup
Ignore file permissions.deny Web UI (Biz/Ent) .cursorignore .clineignore ❌ (sandbox only) .opencodeignore (WIP) .geminiignore ❌
Init command /init /init, /create-instruction Cmd+Shift+P /newrule codex login + web UI /init First-launch wizard ⚠ GUI: Generate Steering Docs
Setup script ❌ copilot-setup-steps.yml βœ… Cloud install script ❌ codex-setup.sh (cloud) ❌ ❌ ❌
Cross-Platform Compatibility
Reads AGENTS.md ❌ (use @import) βœ… βœ… βœ… βœ… βœ… βœ… (v1.20.3) βœ…
Reads other configs ❌ ❌ ❌ βœ… .cursor/ .windsurf/ .claude/ ❌ βœ… CLAUDE.md + .claude/skills/ ❌ ❌ (no CLAUDE.md)
Agent Skills (SKILL.md) βœ… (originator) βœ… .github/skills/ .cursor/skills/ .cline/skills/ βœ… .opencode/skills/ .agent/skills/ βœ… .kiro/skills/

Execution, Coordination & Lifecycle

Feature Claude Code Copilot Cursor Cline Codex OpenCode Antigravity Kiro
Parallelism & Isolation
Parallel execution βœ… Subagents + Teams βœ… Multiple Coding Agents βœ… 8 local + BG agents βœ… use_subagents Read-only; CLI-powered βœ… Cloud tasks βœ… Native parallel tools βœ… Agent Manager βœ… ≀4 subagents
File isolation ⚠ Worktrees βœ… / Teams: LWW βœ… Ephemeral branch βœ… Git worktrees ⚠ Checkpoints βœ… Worktrees + containers Isolation unverified βœ… Optional Git worktree ⚠ Cloud sandbox; local shared
Context isolation βœ… Summary returns βœ… βœ… βœ… βœ… ⚠ Session-scoped + compaction βœ… βœ… Per-subagent context
Communication & Coordination
Comms model βœ… Peer-to-peer Handoffs (VS Code) Parentβ†’child Parentβ†’child βœ… parentβ†’child ❌ No peer comms ❌ ⚠ Hierarchical + summary
Shared task lists βœ… JSON w/ dependencies ❌ ❌ ❌ ❌ ⚠ Todo tool (session) ⚠ Artifacts βœ… Spec list + todo tool
Userβ†’agent access βœ… multi-surface βœ… Issues / PRs / Slack / Teams / Linear / Jira / IDE + 4 more βœ… Web/Slack βœ… CLI + JetBrains + ACP βœ… CLI/IDE/App/Cloud + GH/Slack/Linear βœ… 6 surfaces βœ… Inbox βœ… Mid-turn redirect
Session resumption ⚠ Subagents: βœ… Teams: ❌ βœ… PR iterate + IDE resume βœ… βœ… βœ… resume/fork βœ… SQLite sessions βœ… ⚠ History + switching
Cloud execution ⚠ Web VMs + GH Actions + GitLab CI/CD βœ… GH Actions βœ… AWS VMs ❌ βœ… Codex Cloud ⚠ Headless + GH Actions ❌ βœ… Kiro Web (GA Sep 2026)
Hooks, Review & Quality
Hook events βœ… lifecycle hooks βœ… 8 βœ… 12 βœ… 4 ⚠ 2 experimental experimental βœ… plugin events βœ… Tool / invocation / stop βœ… shell + agent
Code review βœ… /code-review βœ… @copilot + IDE + security ⚠ Bugbot ⚠ Pipe pattern βœ… /review ⚠ GitHub Agent ⚠ Artifacts βœ… Supervised per-hunk
Model comparison ⚠ Manual tiering ⚠ Auto select βœ… Multi-Agent Judging ⚠ ⚠ ❌ ⚠ ❌
Commentable artifacts ❌ ❌ ❌ ❌ ❌ ❌ βœ… 6 types ⚠ Inline hunk chat
3rd-party agents ❌ βœ… Claude + Codex on GitHub + VS Code ❌ ❌ ❌ βœ… Custom agents + MCP + ACP ❌ ⚠ Skills + MCP

Cost & Efficiency

Platform Pricing Approx. Cost Key Optimization
Claude Code Subscription plans + separate API billingas of 2026-08-21 $20 Pro monthly ($17 annual equivalent) Β· Max $100/$200 Β· Team Standard $25 monthly/$20 annual Β· Premium $125 monthly/$100 annual Use the paid-plan allowance for interactive Claude Code; treat Console/API usage as a separate budget
Copilot Plan subscription + AI Creditsas of 2026-08-21 $0 Free Β· $10 Pro/1,500 credits Β· $39 Pro+/7,000 Β· $100 Max/20,000 Β· Business $19/seat Β· Enterprise $39/seat Track AI Credits (1 credit = $0.01), choose plan by monthly allowance, and set organization budgets for overage
Cursor Subscription + two usage pools + on-demandas of 2026-08-21 Start β‚Ή649 India-only Β· Pro $20 Β· Pro Plus $60 Β· Ultra $200 Β· Teams $40/$120 per user Β· Enterprise custom Match plan to Other Models usage ($20/$70/$400 included); use on-demand only when API-rate spend is acceptable
Cline Apache-2.0 clients + BYOK + Cline usage-billing + optional ClinePassas of 2026-08-23 $0 clients; provider/PAYG usage varies; ClinePass $9.99/mo Plan cheap + Act capable. --max-turns. Auto Compact. Per-request cost tracking
Codex ChatGPT plan + optional credits; API separateas of 2026-08-21 Free $0 Β· Go $8 Β· Plus $20 Β· Pro $100/$200 usage tiers Β· Business $20 annual/$25 monthly Β· Enterprise/Edu custom Separate plan allowance, purchased credits and API-token spend; choose the layer that matches interactive vs programmatic use
OpenCode Free (MIT) + OpenCode Black free $0 BYOK β†’ $20–$200 Black Free models available. Ollama: $0 local. Enterprise: custom per-seat $0 core + BYOK (zero markup). Ollama local = truly free. Zen gateway pay-per-token. Free tier models for zero-cost start
Antigravity GA Individual $0 + Google AI plansas of 2026-08-21 $0 Individual Β· $20 Pro Β· $100 Ultra (5x Pro) Β· $200 top Ultra (20x Pro) Start with the weekly-limited Individual tier; choose the 5x or 20x tier only from measured Antigravity usage
Kiro Subscription credits + prepaid add-onsas of 2026-08-21 Free 50 Β· Pro $20/1,000 Β· Pro+ $40/2,000 Β· Pro Max $100/5,000 Β· Power $200/10,000 Β· add-ons $0.04/credit Plan credits expire monthly; prepaid add-ons roll over and expire after 12 months. Enterprise overage is off by default

Strengths & Key Gaps

Claude Code

  • + Peer-to-peer agent messaging (Teams: bidirectional inter-agent comms)
  • + Deep subagent system (3 built-in + custom .md definitions + Teams orchestration)
  • + Broad lifecycle hook coverage; 5 permission modes (allowlist β†’ deny); resumable subagents
+8 more strengths
  • + Terminal/IDE, desktop/web, remote/browser, CI/CD, and SDK access categories
  • + Remote Control: local terminal β†’ mobile/web (no cloud β€” runs on your machine)
  • + Voice Mode: /voice + hold spacebar to talk (current /voice documentation)
  • + Claude Code Security : AI vulnerability scanner (Opus 4.6); 500+ zero-days found in OSS
  • + Skills β†’ subagent delegation via context:fork; Agent SDK (Python + TypeScript) + headless mode
  • + Session teleportation (local ↔ cloud seamless); Auto Memory; Opus 4.6 (1M ctx, 128K output, adaptive thinking)
  • + Open-source sandbox runtime; /compact for long sessions; Plugin marketplace
  • + Opus 4.7 + xhigh effort tier (v2.1.111, Apr 16)
  • - Cloud execution GitHub-only (research preview; no GitLab repos in VMs)
  • - Chrome integration beta (no Bedrock/Vertex); Pro limited to Haiku in Chrome
  • - Teams: no resume; last-write-wins file isolation
+4 more gaps
  • - Claude-only models (no multi-vendor)
  • - No Arena / model A/B comparison
  • - Remote Control: Max-only (Pro coming soon); no Team/Enterprise yet
  • - No GUI IDE (terminal-first; VS Code ext is companion, not fork)

Copilot

  • + Broad IDE + platform reach (VS Code, JetBrains, Xcode, Eclipse, CLI (GA Feb 2026), GitHub.com, Slack, Teams, Linear, Jira (Mar 2026))
  • + VS Code custom-agent handoffs + Agent HQ (mission control for multi-agent coordination)
  • + 3rd-party agents (Claude + Codex assignable on GitHub PRs/Issues)
+6 more strengths
  • + Agentic Memory + built-in code review + security scanning (CodeQL)
  • + Plugins (plugin.json: agents + skills + hooks in one package); 6 hook events
  • + Auto model selection (10% premium discount); 4.7M paid subscribers (as of late 2025)
  • + Coding Agent runs in GitHub Actions (full CI/CD integration)
  • + US + EU data residency + FedRAMP-authorized models (Apr 13) β€” documented enterprise compliance controls
  • + GPT-5.5 GA (Apr 24) and Opus 4.7 GA (Apr 16) on launch day
  • - No peer-to-peer comms or task lists
  • - No Arena / model A/B comparison
  • - Coding Agent costs GH Actions minutes + AI Credits
+3 more gaps
  • - Agent Mode single-session in IDE (no background persistence)
  • - No worktree-based isolation
  • - Custom agents .agent.md: model + handoffs not yet on GitHub.com

Cursor

  • + Subagents (async, nested tree, v2.4+)
  • + Long-Running Agents (multi-hour/day, research preview)
  • + Plugin Marketplace (v2.5: Stripe, Vercel, AWS, etc.)
+8 more strengths
  • + Composer 2 (proprietary thinking model)
  • + Cloud BG Agents (8 parallel); Bugbot PR review (usage billing for new/renewed subscriptions; vendor reports 80% of identified bugs resolved by merge); Automations (Mar 2026: event-triggered cloud agents for Slack/Linear/GitHub/PagerDuty)
  • + .cursor/agents/*.md + Skills + Hooks (12 events) + JetBrains via ACP (Mar 2026)
  • + Cursor Blame (AI attribution, Enterprise)
  • + Multi-Agent Judging; Sandbox network controls
  • + v3.2 (Apr 24): /multitask async subagents + worktrees + multi-root workspaces
  • + Security Review beta + Vulnerability Scanner (Apr 30)
  • + Cursor SDK public beta (Apr 29)
  • - No peer comms or Copilot-style handoffs
  • - Long-running agents research preview only
  • - Closed source IDE (VS Code fork with proprietary additions)
+1 more gaps
  • - No blind model comparison (Arena equivalent)

Cline

  • + Open source (Apache 2.0; 67,384 stars as of 2026-09-03)
  • + gRPC architecture : multi-platform (VS Code, JetBrains, CLI 2.0 (TUI+headless), Zed/Neovim via ACP)
  • + 30+ providers incl. local Ollama; BYOK w/ per-request cost tracking
+5 more strengths
  • + 4 hooks; Plan/Act modes; checkpoints; CLI pipe chaining; Workflows
  • + MCP Marketplace (300+ servers); Skills (always-on since v3.57); Memory Bank
  • + Subagents (v3.58): parallel read-only exploration agents; Teams/Enterprise tiers available
  • + Opus 4.7 (v3.79.0) + GPT-5.5 (v3.82.0) + SAP AI Core + Z AI providers
  • + Enterprise skills management with remote config (v3.80.0)
  • - Subagents currently read-only (research focus, not general-purpose)
  • - No custom agent defs; no agent-to-agent handoffs
  • - No cloud execution; hooks macOS/Linux only (no Windows)
+2 more gaps
  • - No background/headless agents
  • - Extension-based (depends on host IDE, not standalone)

Codex

  • + 5 surfaces (CLI + IDE + App (macOS + Windows) + Cloud) β€” broad Codex-specific reach
  • + Codex App (macOS + Windows Mar 2026): command center with worktrees, Automations (unprompted scheduled work)
  • + Built-in /review + GitHub @codex PR reviews (GPT-5-Codex trained for review)
+5 more strengths
  • + Open-source CLI (Rust, Apache 2.0) + Agent Skills Standard support
  • + Cloud containers + local worktrees; resume & fork sessions
  • + GPT-5.3-Codex-Spark (>1K tok/s on Cerebras); mid-turn steering
  • + Built-in and custom TOML subagents with thread steering, inherited sandboxing, and MCP/skill specialization
  • + v0.128.0 (Apr 30): persistent goal workflows + expanded permission profiles + Bedrock model support
  • - Hooks experimental only (2 events: SessionStart, Stop β€” Mar 2026; full system in design)
  • - Custom-agent TOML format is newer and may evolve; subagent coordination adds token cost
  • - OpenAI models only (no multi-vendor)
+2 more gaps
  • - App: macOS only (Windows alpha)
  • - Local memories are opt-in and off by default

OpenCode

  • + Open-source, provider-agnostic agent (MIT, 203,379β˜… as of 2026-09-03, TypeScript/Bun, hosted and local providers incl. Ollama)
  • + Client/server architecture : Hono HTTP server (port 4096) enables TUI, web, mobile, Docker, IDE control
  • + Deep extensibility : broad documented plugin-event surface, MCP (OAuth 2.0), npm plugins, SKILL.md, custom tools, ACP
+5 more strengths
  • + 6 surfaces : TUI, CLI, web, desktop (Tauri), IDE extensions (VS Code/Cursor/Zed/Windsurf), ACP server
  • + Adaptive thinking (v1.2.7): Claude Sonnet 4.6 + Gemini 3.1 medium reasoning; Ctrl+T toggle
  • + CLAUDE.md compatible (AGENTS.md); GitHub Agent (Actions-based); OpenTUI (Zig + SolidJS, 60fps)
  • + 2.5M+ monthly active developers; 10M+ downloads; 700+ contributors
  • + Actively maintained: v1.18.27 (Sep 2026), ~weekly releases (anomalyco/opencode, TS rewrite)
  • - CVE-2026-22812: CVSS 8.8 unauthenticated local HTTP server in versions <1.0.216; local process or visited malicious-site command execution
  • - No native sandboxing ("UX feature, not security boundary")
  • - No peer-to-peer inter-agent comms
+4 more gaps
  • - No dedicated cloud sandbox (headless + GH Actions only)
  • - Browser limited to webfetch/websearch (no DOM, no screenshots)
  • - Team acknowledged being "overwhelmed" with growth pace
  • - Benchmarks: slower than Claude Code (16m vs 9m, but more thorough)

Antigravity

  • + Agent Manager (Mission Control, up to 5 parallel agents)
  • + Commentable artifacts (6 types: plans, diffs, walkthroughs, screenshots, browser recordings)
  • + Browser subagent w/ WebP video recording (dedicated Gemini 2.5 CU model)
+6 more strengths
  • + Plan-dependent roster: Gemini 3.8/3.7/3.6 Flash, Gemini 3.1 Pro, Claude Sonnet/Opus 4.6, GPT-OSS-120b
  • + GA Individual $0; Google AI Pro $20; Ultra $100 (5x) / $200 (20x)
  • + Gemini 3.1 Pro (Feb 19, 2026): 80.6% SWE-bench, 77.1% ARC-AGI-2, 94.3% GPQA Diamond; customtools endpoint
  • + Lifecycle hooks via hooks.json
  • + Custom Markdown subagents
  • + Optional branch-mode Git worktree isolation
  • - No background/headless agents (foreground IDE only)
  • - Closed source; code transmitted to Google data centers
  • - Security bundle remains withheld pending primary advisory locators

Kiro

  • + Spec-driven development as the core unit of work (Requirementsβ†’Designβ†’Tasks; the spec is the source of truth, code is a build artifact)
  • + Event-driven agent hooks fire on save/PR/repo events (run tests, update docs, cascade spec changes)
  • + Steering files (.kiro/steering/) for persistent project knowledge with 4 inclusion modes
+3 more strengths
  • + Powers extension system (launch partners Figma + Netlify) + Kiro CLI sharing the same steering/MCP config
  • + Property-based testing measures whether code matches the spec
  • + AWS-grade trust: HIPAA-eligible, GovCloud, IP indemnity, no training on paid-tier content
  • - Kiro autonomous agents remain preview while Kiro Web is GA (autonomous agent is one-per-developer)
  • - No bring-your-own-key or local-model routing; models are Kiro-hosted, and current docs list OpenAI, Anthropic, and selected open-weight families but not Gemini
  • - Credits do not roll over month-to-month (penalizes bursty usage)
+1 more gaps
  • - Notable first-year security track record: multiple workspace-trust code-execution CVEs in 2026

Known Risks & Mitigations (Feb 2026)

Cursor DuneSlide: two critical sandbox escapes fixed in 3.0criticalconfirmed

CVE-2026-50548 lets an agent-controlled working directory expose writable paths outside the workspace; CVE-2026-50549 abuses symlink/path-canonicalization fallback. Both affect Cursor <3.0 and can enable unsandboxed RCE without interaction beyond a benign prompt. NVD: CVSS v3.1 9.8; CNA: CVSS v4 9.3.
β†’ Mitigation: Update to Cursor 3.0 or later; both vendor advisories list 3.0 as the patched version.
Bound assertion: β€œCVE-2026-50548”Due soonclaim-cursor-duneslide-cve Β· Cursor advisory CVE-2026-50548 (official Β· official_updates) Β· recorded verification 2026-09-03 Β· review due 2026-09-17

Cursor BG Agents auto-run terminal commandshighobserved

Background mode executes shell commands without per-command approval. Data retained ~a few days on remote VMs.
β†’ Mitigation: Use interactive mode for sensitive repos. Review BG agent PRs before merging.
Unbound legacy projection Β· Legacy background-agent risk has not yet been atomized into a claim.

OpenCode advisory GHSA-vxw4-wv6m-9hhh Β· CVSS 8.8highconfirmed

opencode-ai versions below 1.0.216 are affected; 1.0.216 is patched. The running local HTTP server could execute commands with the user's privileges. The advisory confirms a malicious-site vector in Firefox and says user interaction is required.
β†’ Named criterion: affected versions fail; an unknown selected version is unconfirmed; 1.0.216 lies outside this advisory’s affected range. This does not establish general product safety.
Bound assertion: β€œopencode-ai versions below 1.0.216 are affected; 1.0.216 is patched”Currentclaim-opencode-advisory-version-range Β· OpenCode maintainer advisory GHSA-vxw4-wv6m-9hhh (official Β· official_updates) Β· recorded verification 2026-09-09 Β· review due 2026-10-09
Bound assertion: β€œmalicious-site vector in Firefox and says user interaction is required”Currentclaim-opencode-advisory-browser-vector Β· OpenCode maintainer advisory GHSA-vxw4-wv6m-9hhh (official Β· official_updates) Β· recorded verification 2026-09-09 Β· review due 2026-10-09
Bound assertion: β€œCVSS 8.8”Currentclaim-opencode-advisory-cvss Β· OpenCode maintainer advisory GHSA-vxw4-wv6m-9hhh (official Β· official_updates) Β· recorded verification 2026-09-09 Β· review due 2026-10-09

Cross-platform: rules-file prompt injection (all 8 platforms)highobserved

All platforms load instruction files from workspace (CLAUDE.md, copilot-instructions.md, .cursorrules, .clinerules, AGENTS.md, GEMINI.md, Kiro .kiro/steering/). Malicious repos can include crafted rules that exfiltrate context, override safety behaviors, or inject hidden instructions. IDEsaster research (Dec 2025) demonstrated 24 CVEs across 30+ IDE extensions including configuration manipulation and data leakage via tool-use loops.
β†’ Mitigation: Audit rules files in cloned repos before opening. Use .gitignore for sensitive files. Enable permission prompts (avoid Yolo/Turbo/Auto-approve modes on untrusted repos). Enterprise: deploy org-level rules via MDM/policy.
Unbound legacy projection Β· Cross-platform prompt-injection summary has not yet been atomized into a claim.

⚠ Stale evidence β€” Antigravity 2.0 release-quality reports (community-only)highobserved

Evidence status: stale; last verified 2026-06-12, with no official Google acknowledgement. Community threads reported removed SSH and extension support, install/update bugs, context loss during migration from 1.x, credit-allotment changes, and a rollback procedure to 1.23.2. Treat these as dated issue signals, not current platform-wide facts.
β†’ Evidence action: re-verify against current official release notes before deriving version advice; monitor discuss.ai.google.dev only as a community issue signal.
Bound assertion: β€œStale evidence”Withheldclaim-antigravity-2-0-release-quality Β· Forum: 'Antigravity 2.0 is awful, here's how to get the previous version' (secondary Β· manual_research) Β· recorded verification 2026-06-12 Β· review due 2026-07-12

Kiro workspace-trust code-execution CVEs (2026 cluster)highconfirmed

Three 2026 workspace-related code-execution entries in the GitHub Advisory Database, all marked Unreviewed: CVE-2026-5429 (CVSS v4 7.1) fixed in 0.8.140; CVE-2026-4295 (CVSS v4 8.5) fixed in 0.8.0; CVE-2026-0830 (CVSS v4 8.4) fixed in 0.6.18. AWS bulletins 2026-012, 2026-009, and 2026-001 confirm the respective version boundaries.
β†’ Mitigation: Keep Kiro updated (recurring fix pattern); never open or trust untrusted workspaces; in enterprise restrict via IAM Identity Center plus MCP/extension allowlists.
Bound assertion: β€œCVE-2026-5429”Currentclaim-kiro-cve-cluster Β· AWS security bulletin 2026-009 (official Β· official_updates) Β· recorded verification 2026-08-23 Β· review due 2026-11-23

Cross-platform: HalluSquatting adversarial package/skill hallucinationhighobserved

The arXiv:2607.07433 preprint (Jul 8, 2026) evaluates nine application variants, including six coding assistants. It reports hallucinated-resource generation up to 85% for repository cloning and 100% for skill installation; end-to-end tool/RCE success is a separate 20-65% for coding assistants and 40-100% for personal assistants.
β†’ Mitigation: Pin and verify dependencies + skills before install; disable auto-install of suggested packages/skills on untrusted repos; enforce lockfiles + provenance checks.
Sources: ref arXiv:2607.07433
Bound assertion: β€œHalluSquatting”Due soonclaim-hallusquatting-cross-platform Β· Beware of Agentic Botnets (arXiv preprint) (secondary Β· manual_research) Β· recorded verification 2026-09-03 Β· review due 2026-09-17

Copilot Coding Agent costs combine Actions and AI Creditsmediumconfirmed

Coding Agent consumes GitHub Actions minutes and AI Credits. One AI Credit equals $0.01; model token rates determine consumption.
β†’ Mitigation: Track the included credit pool, set organization budgets, and reserve Coding Agent for well-defined issues.
Unbound legacy projection Β· Legacy cost-risk prose has a source link but no atomic claim binding for this card.

Kiro autonomous-agent blast radius (contested)mediumobserved

Dec 2025 reports (Financial Times) described an internal Amazon use of Kiro that deleted and recreated a production environment, taking AWS Cost Explorer offline in one region for ~13 hours. Amazon disputes the AI framing, attributing the incident to a misconfigured, over-permissioned access-control role rather than the model, and calls it a limited single-service event in one of 39 regions. The event is corroborated; AI causation is contested.
β†’ Mitigation: Require human-approval gates and least-privilege IAM roles for any autonomous or destructive agent actions; the documented root cause was an over-permissioned role, not the model.
Unbound legacy projection Β· Contested autonomous-agent incident remains outside the current atomic claim set.

Copilot custom agent handoffs are IDE-onlylowconfirmed

The model and handoffs properties in .agent.md work in VS Code and IDE agents, but are not yet supported for Coding Agent on GitHub.com.
β†’ Mitigation: Design agent chains for IDE use. For Coding Agent, use single-purpose custom agents.
Bound assertion: β€œcustom agent handoffs are IDE-only”Currentclaim-copilot-agent-handoffs Β· VS Code custom agents (official Β· official_docs) Β· recorded verification 2026-08-23 Β· review due 2026-09-23

Cursor plans don't auto-persistlowobserved

Plans only save when user explicitly clicks "Save to workspace." Saved to .cursor/plans/.
β†’ Mitigation: Always Save-to-workspace before modifying plans.
Sources: forum forum.cursor.com (user reports)
Unbound legacy projection Β· Community-reported persistence limitation has no current atomic claim.

Cursor security detail withheld pending a primary advisorylowtheoretical

A previous report version asserted affected-vector and patched-version details for CVE-2026-26268. The current audit could not locate a stable primary advisory, so those details are not presented as verified facts.
β†’ Evidence action: keep this item stale and do not derive a version recommendation until the primary advisory is captured and reviewed.
Bound assertion: β€œsecurity detail withheld pending a primary advisory”Withheldclaim-cursor-cve-sandbox Β· Full platform source audit β€” evidence withheld (secondary Β· manual_research) Β· recorded verification 2026-05-04 Β· review due 2026-06-03

Claude Code Teams: no resume, last-write-winslowobserved

Agent Teams sessions can't be resumed. Multiple teammates on same file risk overwrites.
β†’ Mitigation: Assign non-overlapping file scopes. Use subagents (which do resume) for smaller tasks.
Unbound legacy projection Β· Compound team-isolation limitations have not yet been atomized into a claim.

Claude Code Web VMs: GitHub-only, research previewlowconfirmed

Cloud execution runs on Anthropic-managed VMs but currently supports GitHub repos only β€” no GitLab/Bitbucket. Research preview status. Chrome integration requires claude.ai account (not available via Bedrock/Vertex/Foundry).
β†’ Mitigation: Use GitHub Actions integration for CI/CD. For non-GitHub repos, run locally or via SSH.
Unbound legacy projection Β· Compound cloud-surface limitations have not yet been atomized into a claim.

Cline subagents are read-onlylowconfirmed

Research subagents cannot write files, use browser, or call MCP tools.
β†’ Mitigation: Use subagents for research, apply findings in main agent loop.
Unbound legacy projection Β· Subagent limitation has a source but no dedicated atomic claim binding.

Antigravity security detail withheld pending primary advisorieslowtheoretical

A previous report version bundled several security and data-residency allegations. The underlying local evidence snapshot is missing and stable primary advisory locators were not found, so the allegations are not presented as verified facts.
β†’ Evidence action: retain the stale record for audit history, but withhold operational conclusions until exact primary advisories are captured and reviewed.
Bound assertion: β€œsecurity detail withheld pending primary advisories”Withheldclaim-antigravity-security Β· Full platform source audit β€” evidence withheld (secondary Β· manual_research) Β· recorded verification 2026-06-12 Β· review due 2026-07-12

Harness Engineering

Research baseline

Curated recommendations with assertion-level provenance and a machine-checked lifecycle contract. This is not a complete execution runtime or proof of comparative effectiveness.

6 practices Β· 18 assertions Β· 6 cited sources

Greenfield

Curated hypothesis
  1. Create the smallest context and verification mapplan Β· Small root map and progressive context, Executable verification inside the work loop
  2. Build one observable vertical sliceexecute Β· Runtime and process evidence, Executable verification inside the work loop
  3. Promote repeated failures only after evidenceretro Β· Promote recurring failures into narrow guardrails

Brownfield

Curated hypothesis
  1. Persist observed state and existing verification debtinspect Β· Durable state and resumable handoffs, Runtime and process evidence
  2. Bound blast radius and irreversible actionsapprove Β· Bounded actions and human authority
  3. Prove no new regression at affected boundariesverify Β· Executable verification inside the work loop

Runbook map

  1. Evidence MapSource and evidence inventory
  2. Universal CoreDurable state and resumable handoffs, Small root map and progressive context, Executable verification inside the work loop, Bounded actions and human authority, Runtime and process evidence, Promote recurring failures into narrow guardrails
  3. Greenfield ProfileSmall root map and progressive context, Executable verification inside the work loop, Runtime and process evidence
  4. Brownfield ProfileDurable state and resumable handoffs, Executable verification inside the work loop, Bounded actions and human authority, Runtime and process evidence
  5. Failure and Retro LoopPromote recurring failures into narrow guardrails
  6. Practice CardsDurable state and resumable handoffs, Small root map and progressive context, Executable verification inside the work loop, Bounded actions and human authority, Runtime and process evidence, Promote recurring failures into narrow guardrails
  7. Contested and Research PendingSmall root map and progressive context, Bounded actions and human authority, Promote recurring failures into narrow guardrails

Practice cards

Durable state and resumable handoffsuniversal
RecommendationSupportedActive

Persist task state, decisions, verification evidence, and the next action outside chat so a fresh session can resume from repository-visible artifacts.

Qualification: Directly described in an inspected first-party implementation report.

  • Effective harnesses for long-running agents β€” The long-running agent problem and two-fold solution Β· implements / direct / accepted Β· full inspection Β· active Β· review due 2026-11-23Limit: First-party report about Anthropic's own long-running web-app harness; it does not establish universal effectiveness.
ApplicabilitySupportedActive

Apply this pattern when work can cross a context reset, human handoff, interruption, or multi-session boundary.

Qualification: The cited report is specifically scoped to multi-session work.

  • Effective harnesses for long-running agents β€” Getting up to speed Β· supports / direct / accepted Β· full inspection Β· active Β· review due 2026-11-23Limit: First-party report about Anthropic's own long-running web-app harness; it does not establish universal effectiveness.
LimitationSupportedActive

The evidence supports durable handoff artifacts, not one universal filename, schema, or storage backend.

Qualification: This limitation prevents copying one vendor's implementation as a standard.

  • Effective harnesses for long-running agents β€” Environment management Β· supports / inferred / limited Β· full inspection Β· active Β· review due 2026-11-23Limit: First-party report about Anthropic's own long-running web-app harness; it does not establish universal effectiveness.
Small root map and progressive contextuniversal
RecommendationSupportedActive

Keep the always-loaded instruction surface concise and route agents to scoped context on demand.

Qualification: Directly recommended by an inspected first-party field report and an independently maintained course.

  • Learn Harness Engineering β€” Lecture 04: progressive disclosure Β· supports / direct / limited Β· full inspection Β· active Β· review due 2026-11-20Limit: Course and pattern catalog, not peer-reviewed evidence; bundled validators and examples have documented internal gaps.
  • Harness engineering: leveraging Codex in an agent-first world β€” Repository knowledge as the system of record Β· implements / direct / accepted Β· full inspection Β· active Β· review due 2026-11-23Limit: First-party account of OpenAI's own internal repository; it reports operating patterns but does not isolate their causal effect.
ApplicabilitySupportedActive

Use progressive context when a repository has several domains, workflows, or provider-specific instruction files.

Qualification: The field report, open instruction-file specification, and inspected course agree on scoped routing for multi-domain repositories.

  • Learn Harness Engineering β€” Lectures 03–04 Β· supports / inferred / limited Β· full inspection Β· active Β· review due 2026-11-20Limit: Course and pattern catalog, not peer-reviewed evidence; bundled validators and examples have documented internal gaps.
  • Harness engineering: leveraging Codex in an agent-first world β€” A short AGENTS.md as a table of contents Β· implements / direct / accepted Β· full inspection Β· active Β· review due 2026-11-23Limit: First-party account of OpenAI's own internal repository; it reports operating patterns but does not isolate their causal effect.
  • AGENTS.md β€” a simple, open format for guiding coding agents β€” Nested AGENTS.md and closest-file precedence Β· defines / direct / accepted Β· full inspection Β· active Β· review due 2026-11-23Limit: Interoperability convention for instruction discovery and scoping; it does not measure task-quality outcomes.
LimitationCurated HypothesisActive

No universal line or token limit is established by the inspected evidence.

Qualification: Quantitative instruction-budget claims are intentionally excluded.

  • Learn Harness Engineering β€” Lecture 04 examples Β· supports / inferred / limited Β· full inspection Β· active Β· review due 2026-11-20Limit: Course and pattern catalog, not peer-reviewed evidence; bundled validators and examples have documented internal gaps.
Executable verification inside the work loopuniversal
RecommendationSupportedActive

Define executable checks before completion and keep termination judgment outside the implementer's self-report.

Qualification: Directly supported by inspected first-party reports and course material.

  • Effective harnesses for long-running agents β€” Testing Β· implements / direct / accepted Β· full inspection Β· active Β· review due 2026-11-23Limit: First-party report about Anthropic's own long-running web-app harness; it does not establish universal effectiveness.
  • Learn Harness Engineering β€” Lectures 09–10 Β· supports / direct / accepted Β· full inspection Β· active Β· review due 2026-11-20Limit: Course and pattern catalog, not peer-reviewed evidence; bundled validators and examples have documented internal gaps.
  • Demystifying evals for AI agents β€” Task, trial, grader, trace, outcome, and eval-harness definitions Β· defines / direct / accepted Β· full inspection Β· active Β· review due 2026-11-23Limit: Methodology from a model vendor; examples and recommendations do not by themselves prove a harness-specific causal effect.
ApplicabilitySupportedActive

Select checks by affected boundary and risk, including runtime or end-to-end evidence when user journeys change.

Qualification: The sources describe testing and evaluator feedback at application boundaries.

  • Effective harnesses for long-running agents β€” Testing Β· implements / direct / accepted Β· full inspection Β· active Β· review due 2026-11-23Limit: First-party report about Anthropic's own long-running web-app harness; it does not establish universal effectiveness.
  • Harness design for long-running application development β€” The architecture Β· implements / direct / accepted Β· full inspection Β· active Β· review due 2026-11-23Limit: Vendor experiment; planner/generator/evaluator results do not isolate every harness component.
  • Demystifying evals for AI agents β€” Stable coding environments and deterministic graders Β· supports / direct / accepted Β· full inspection Β· active Β· review due 2026-11-23Limit: Methodology from a model vendor; examples and recommendations do not by themselves prove a harness-specific causal effect.
LimitationCurated HypothesisActive

A green check or skipped test does not prove correctness outside the behavior actually exercised.

Qualification: This course-derived scope limitation is not an effectiveness claim.

  • Learn Harness Engineering β€” Harness validator limitations Β· supports / inferred / limited Β· full inspection Β· active Β· review due 2026-11-20Limit: Course and pattern catalog, not peer-reviewed evidence; bundled validators and examples have documented internal gaps.
  • Demystifying evals for AI agents β€” Grade outcomes rather than prescribed paths Β· supports / direct / limited Β· full inspection Β· active Β· review due 2026-11-23Limit: Methodology from a model vendor; examples and recommendations do not by themselves prove a harness-specific causal effect.
Bounded actions and human authorityuniversal
RecommendationCurated HypothesisActive

Bound retries and side effects, and require human authority at irreversible, semantic, or production boundaries.

Qualification: Risk-control recommendation synthesized from the inspected course; no universal optimum is claimed.

  • Learn Harness Engineering β€” Lecture 13: bounded autonomous loops Β· supports / direct / limited Β· full inspection Β· active Β· review due 2026-11-20Limit: Course and pattern catalog, not peer-reviewed evidence; bundled validators and examples have documented internal gaps.
ApplicabilityCurated HypothesisActive

Choose approval and isolation controls from blast radius, reversibility, data sensitivity, and external side effects.

Qualification: Thresholds remain project-specific.

  • Learn Harness Engineering β€” Lifecycle and tool registry references Β· supports / inferred / limited Β· full inspection Β· active Β· review due 2026-11-20Limit: Course and pattern catalog, not peer-reviewed evidence; bundled validators and examples have documented internal gaps.
LimitationCurated HypothesisActive

Human review at every step can become a bottleneck; gates should concentrate on material boundaries.

Qualification: The evidence does not establish one portable approval threshold.

  • Learn Harness Engineering β€” Lecture 13 loop safeguards Β· supports / inferred / limited Β· full inspection Β· active Β· review due 2026-11-20Limit: Course and pattern catalog, not peer-reviewed evidence; bundled validators and examples have documented internal gaps.
Runtime and process evidenceuniversal
RecommendationSupportedActive

Record bounded runtime outputs and process evidence so failures can be attributed to a check, input, decision, or transition.

Qualification: Structured artifacts and evaluator feedback are directly described in the inspected reports.

  • Harness design for long-running application development β€” The architecture Β· implements / direct / accepted Β· full inspection Β· active Β· review due 2026-11-23Limit: Vendor experiment; planner/generator/evaluator results do not isolate every harness component.
  • Demystifying evals for AI agents β€” Trace and outcome as distinct evaluation evidence Β· defines / direct / accepted Β· full inspection Β· active Β· review due 2026-11-23Limit: Methodology from a model vendor; examples and recommendations do not by themselves prove a harness-specific causal effect.
ApplicabilitySupportedActive

Capture process evidence when work is long-running, delegated, expensive, or difficult to reproduce from the final diff alone.

Qualification: The cited report is scoped to long-running delegated application development.

  • Harness design for long-running application development β€” Scaling to full-stack coding Β· supports / direct / accepted Β· full inspection Β· active Β· review due 2026-11-23Limit: Vendor experiment; planner/generator/evaluator results do not isolate every harness component.
LimitationSupportedActive

Logs and traces explain execution but do not themselves establish product correctness or factual grounding.

Qualification: Rendered as a material limitation beside the recommendation.

  • Harness design for long-running application development β€” Evaluator limitations and iteration Β· supports / inferred / limited Β· full inspection Β· active Β· review due 2026-11-23Limit: Vendor experiment; planner/generator/evaluator results do not isolate every harness component.
Promote recurring failures into narrow guardrailsuniversal
RecommendationSupportedActive

Convert recurring or escaped failures into the narrowest reproducible test, validator, policy check, or eval that detects the defect.

Qualification: An inspected first-party field report and an independent course both recommend turning recurring failures into maintained mechanical checks.

  • Learn Harness Engineering β€” Lecture 10: repeated review feedback becomes a check Β· supports / direct / limited Β· full inspection Β· active Β· review due 2026-11-20Limit: Course and pattern catalog, not peer-reviewed evidence; bundled validators and examples have documented internal gaps.
  • Harness engineering: leveraging Codex in an agent-first world β€” Recurring cleanup work and mechanical enforcement Β· implements / direct / accepted Β· full inspection Β· active Β· review due 2026-11-23Limit: First-party account of OpenAI's own internal repository; it reports operating patterns but does not isolate their causal effect.
ApplicabilitySupportedActive

Apply this promotion after repeated failures, escaped defects, or review feedback that recurs across tasks.

Qualification: The sources converge on recurring drift or repeated review findings as the trigger for a narrow guardrail.

  • Learn Harness Engineering β€” Lectures 10 and 12 Β· supports / inferred / limited Β· full inspection Β· active Β· review due 2026-11-20Limit: Course and pattern catalog, not peer-reviewed evidence; bundled validators and examples have documented internal gaps.
  • Harness engineering: leveraging Codex in an agent-first world β€” Recurring cleanup work Β· implements / direct / accepted Β· full inspection Β· active Β· review due 2026-11-23Limit: First-party account of OpenAI's own internal repository; it reports operating patterns but does not isolate their causal effect.
LimitationCurated HypothesisActive

One-off mistakes should not automatically add permanent root-prompt prose or an expensive global gate.

Qualification: This guards against harness bloat and Goodhart pressure.

  • Learn Harness Engineering β€” Lecture 12: cleanup and ablation Β· supports / inferred / limited Β· full inspection Β· active Β· review due 2026-11-20Limit: Course and pattern catalog, not peer-reviewed evidence; bundled validators and examples have documented internal gaps.

Project Initialization β€” Manual Prompt Playbook

Manual init beats /init by ~7 percentage points. Human-curated context files yield ~+4 percentage-point improvement on agent task success vs none; auto-generated files measure ~-3 pp vs none. Frontier LLMs reliably follow ~150-200 instructions; the system prompt alone consumes ~50 of that budget. Keep context files under 300 lines, ideally under 60. Use the universal prompt below; then run it with the per-platform tail line that writes to the right file in the right format.

Universal 3-phase init prompt

You are bootstrapping a context file for AI coding agents. Do NOT auto-write it. Do this in three phases.

# PHASE 1 β€” SCAN (read-only, no edits)

1. List top-level structure: run `ls -la` then `find . -maxdepth 3 -type d -not -path '*/.*' | sort`.
2. Detect stack: read whichever of these exist β€” package.json, pyproject.toml, requirements.txt, Cargo.toml, go.mod, Gemfile, pom.xml, build.gradle, deno.json, *.csproj.
3. Read existing docs: README*, CONTRIBUTING*, ARCHITECTURE*, docs/index*, .editorconfig.
4. Detect vendor instruction files and per-vendor config already present: CLAUDE.md, AGENTS.md, GEMINI.md, .cursor/rules/*.mdc, .github/copilot-instructions.md, .windsurf/rules/*.md, .clinerules (or .clinerules/ directory), opencode.json.
5. Identify the most-edited files: `git log --since='90 days ago' --pretty=format: --name-only | sort | uniq -c | sort -rn | head -25`.
6. Run `--help` (or read help docs) for any non-obvious tool you find in scripts/Makefile (e.g. `make help`, `bun run`, custom CLI). Do NOT execute build/test commands.

# PHASE 2 β€” DRAFT (write a draft, NOT the final file)

Produce a draft with EXACTLY these sections, in this order. Keep total under 300 lines; target 60.

**WHAT** β€” Stack, languages, frameworks, package manager, runtime versions. One line per item. Only versions you actually verified.

**WHY** β€” Two sentences: what this project is for, and the one non-obvious thing about how it is shaped (monorepo? plugin host? CLI + lib? data pipeline?).

**HOW** β€” Bullet list of exact commands. Install, lint, typecheck, test (unit + integration if separate), build, run dev. Use the actual package manager (bun vs pnpm vs npm vs uv vs poetry vs cargo), not a generic placeholder.

**NON-OBVIOUS** β€” Anything an experienced engineer would not guess from reading the code: required env vars and where to get them, services that must run locally, secrets handling, custom DSLs, generated code that should never be edited by hand, code paths that look dead but are entry points.

**BOUNDARIES** β€” Three tiers in this exact format:
  - βœ… Always OK to modify: <paths/patterns>
  - ⚠️ Ask first: <paths/patterns>
  - 🚫 Never touch: <paths/patterns>

**KEY FILES** β€” Pointers (not copies) to deeper docs and the few files that define the architecture. Use relative paths.

Write the draft to a temp path so I can review before it replaces the live file: print it to stdout AND save to `./.init-draft.md`.

# PHASE 3 β€” CURATE (I do this, you wait)

Stop after Phase 2. Do not write to <<OUTPUT_FILE>>. Wait for me to review the draft, strip anything a linter or formatter already enforces, replace any inline code snippets with pointers to source files, and only then I will copy the curated content into <<OUTPUT_FILE>>.

For each platform card below, append the platform-specific tail line. The tail tells the agent the exact output path + format quirks.

Claude Code

CLAUDE.md
Project root. 4-level hierarchy: Enterprise managed β†’ User (~/.claude/CLAUDE.md) β†’ Project (CLAUDE.md) β†’ subdirectories. Project-level is more specific and wins on conflicts.
You are bootstrapping a context file for AI coding agents. Do NOT auto-write it. Do this in three phases.

# PHASE 1 β€” SCAN (read-only, no edits)

1. List top-level structure: run `ls -la` then `find . -maxdepth 3 -type d -not -path '*/.*' | sort`.
2. Detect stack: read whichever of these exist β€” package.json, pyproject.toml, requirements.txt, Cargo.toml, go.mod, Gemfile, pom.xml, build.gradle, deno.json, *.csproj.
3. Read existing docs: README*, CONTRIBUTING*, ARCHITECTURE*, docs/index*, .editorconfig.
4. Detect vendor instruction files and per-vendor config already present: CLAUDE.md, AGENTS.md, GEMINI.md, .cursor/rules/*.mdc, .github/copilot-instructions.md, .windsurf/rules/*.md, .clinerules (or .clinerules/ directory), opencode.json.
5. Identify the most-edited files: `git log --since='90 days ago' --pretty=format: --name-only | sort | uniq -c | sort -rn | head -25`.
6. Run `--help` (or read help docs) for any non-obvious tool you find in scripts/Makefile (e.g. `make help`, `bun run`, custom CLI). Do NOT execute build/test commands.

# PHASE 2 β€” DRAFT (write a draft, NOT the final file)

Produce a draft with EXACTLY these sections, in this order. Keep total under 300 lines; target 60.

**WHAT** β€” Stack, languages, frameworks, package manager, runtime versions. One line per item. Only versions you actually verified.

**WHY** β€” Two sentences: what this project is for, and the one non-obvious thing about how it is shaped (monorepo? plugin host? CLI + lib? data pipeline?).

**HOW** β€” Bullet list of exact commands. Install, lint, typecheck, test (unit + integration if separate), build, run dev. Use the actual package manager (bun vs pnpm vs npm vs uv vs poetry vs cargo), not a generic placeholder.

**NON-OBVIOUS** β€” Anything an experienced engineer would not guess from reading the code: required env vars and where to get them, services that must run locally, secrets handling, custom DSLs, generated code that should never be edited by hand, code paths that look dead but are entry points.

**BOUNDARIES** β€” Three tiers in this exact format:
  - βœ… Always OK to modify: <paths/patterns>
  - ⚠️ Ask first: <paths/patterns>
  - 🚫 Never touch: <paths/patterns>

**KEY FILES** β€” Pointers (not copies) to deeper docs and the few files that define the architecture. Use relative paths.

Write the draft to a temp path so I can review before it replaces the live file: print it to stdout AND save to `./.init-draft.md`.

# PHASE 3 β€” CURATE (I do this, you wait)

Stop after Phase 2. Do not write to <<OUTPUT_FILE>>. Wait for me to review the draft, strip anything a linter or formatter already enforces, replace any inline code snippets with pointers to source files, and only then I will copy the curated content into <<OUTPUT_FILE>>.

# PHASE 3 TAIL FOR Claude Code
When I confirm, write the curated content to ./CLAUDE.md at the repository root.
Format quirks: Plain Markdown. Use `@<path>` to import other files (the standard way to get AGENTS.md interop until issue #6235 lands). Tooltip: `@AGENTS.md` is the canonical workaround.

GitHub Copilot

.github/copilot-instructions.md
Repo-level instructions. For file-scoped rules also create `.github/instructions/*.instructions.md` with YAML frontmatter `applyTo` glob. Also natively reads AGENTS.md.
You are bootstrapping a context file for AI coding agents. Do NOT auto-write it. Do this in three phases.

# PHASE 1 β€” SCAN (read-only, no edits)

1. List top-level structure: run `ls -la` then `find . -maxdepth 3 -type d -not -path '*/.*' | sort`.
2. Detect stack: read whichever of these exist β€” package.json, pyproject.toml, requirements.txt, Cargo.toml, go.mod, Gemfile, pom.xml, build.gradle, deno.json, *.csproj.
3. Read existing docs: README*, CONTRIBUTING*, ARCHITECTURE*, docs/index*, .editorconfig.
4. Detect vendor instruction files and per-vendor config already present: CLAUDE.md, AGENTS.md, GEMINI.md, .cursor/rules/*.mdc, .github/copilot-instructions.md, .windsurf/rules/*.md, .clinerules (or .clinerules/ directory), opencode.json.
5. Identify the most-edited files: `git log --since='90 days ago' --pretty=format: --name-only | sort | uniq -c | sort -rn | head -25`.
6. Run `--help` (or read help docs) for any non-obvious tool you find in scripts/Makefile (e.g. `make help`, `bun run`, custom CLI). Do NOT execute build/test commands.

# PHASE 2 β€” DRAFT (write a draft, NOT the final file)

Produce a draft with EXACTLY these sections, in this order. Keep total under 300 lines; target 60.

**WHAT** β€” Stack, languages, frameworks, package manager, runtime versions. One line per item. Only versions you actually verified.

**WHY** β€” Two sentences: what this project is for, and the one non-obvious thing about how it is shaped (monorepo? plugin host? CLI + lib? data pipeline?).

**HOW** β€” Bullet list of exact commands. Install, lint, typecheck, test (unit + integration if separate), build, run dev. Use the actual package manager (bun vs pnpm vs npm vs uv vs poetry vs cargo), not a generic placeholder.

**NON-OBVIOUS** β€” Anything an experienced engineer would not guess from reading the code: required env vars and where to get them, services that must run locally, secrets handling, custom DSLs, generated code that should never be edited by hand, code paths that look dead but are entry points.

**BOUNDARIES** β€” Three tiers in this exact format:
  - βœ… Always OK to modify: <paths/patterns>
  - ⚠️ Ask first: <paths/patterns>
  - 🚫 Never touch: <paths/patterns>

**KEY FILES** β€” Pointers (not copies) to deeper docs and the few files that define the architecture. Use relative paths.

Write the draft to a temp path so I can review before it replaces the live file: print it to stdout AND save to `./.init-draft.md`.

# PHASE 3 β€” CURATE (I do this, you wait)

Stop after Phase 2. Do not write to <<OUTPUT_FILE>>. Wait for me to review the draft, strip anything a linter or formatter already enforces, replace any inline code snippets with pointers to source files, and only then I will copy the curated content into <<OUTPUT_FILE>>.

# PHASE 3 TAIL FOR GitHub Copilot
When I confirm, write the curated content to ./.github/copilot-instructions.md.
Format quirks: Plain Markdown. For Coding Agent (web/PR) also create `.github/workflows/copilot-setup-steps.yml` declaring runtime + deps + lint/test. That file IS the brownfield runner spec.

Cursor

.cursor/rules/project.mdc
Modern .cursor/rules/*.mdc directory. Legacy .cursorrules at root still read but discouraged.
You are bootstrapping a context file for AI coding agents. Do NOT auto-write it. Do this in three phases.

# PHASE 1 β€” SCAN (read-only, no edits)

1. List top-level structure: run `ls -la` then `find . -maxdepth 3 -type d -not -path '*/.*' | sort`.
2. Detect stack: read whichever of these exist β€” package.json, pyproject.toml, requirements.txt, Cargo.toml, go.mod, Gemfile, pom.xml, build.gradle, deno.json, *.csproj.
3. Read existing docs: README*, CONTRIBUTING*, ARCHITECTURE*, docs/index*, .editorconfig.
4. Detect vendor instruction files and per-vendor config already present: CLAUDE.md, AGENTS.md, GEMINI.md, .cursor/rules/*.mdc, .github/copilot-instructions.md, .windsurf/rules/*.md, .clinerules (or .clinerules/ directory), opencode.json.
5. Identify the most-edited files: `git log --since='90 days ago' --pretty=format: --name-only | sort | uniq -c | sort -rn | head -25`.
6. Run `--help` (or read help docs) for any non-obvious tool you find in scripts/Makefile (e.g. `make help`, `bun run`, custom CLI). Do NOT execute build/test commands.

# PHASE 2 β€” DRAFT (write a draft, NOT the final file)

Produce a draft with EXACTLY these sections, in this order. Keep total under 300 lines; target 60.

**WHAT** β€” Stack, languages, frameworks, package manager, runtime versions. One line per item. Only versions you actually verified.

**WHY** β€” Two sentences: what this project is for, and the one non-obvious thing about how it is shaped (monorepo? plugin host? CLI + lib? data pipeline?).

**HOW** β€” Bullet list of exact commands. Install, lint, typecheck, test (unit + integration if separate), build, run dev. Use the actual package manager (bun vs pnpm vs npm vs uv vs poetry vs cargo), not a generic placeholder.

**NON-OBVIOUS** β€” Anything an experienced engineer would not guess from reading the code: required env vars and where to get them, services that must run locally, secrets handling, custom DSLs, generated code that should never be edited by hand, code paths that look dead but are entry points.

**BOUNDARIES** β€” Three tiers in this exact format:
  - βœ… Always OK to modify: <paths/patterns>
  - ⚠️ Ask first: <paths/patterns>
  - 🚫 Never touch: <paths/patterns>

**KEY FILES** β€” Pointers (not copies) to deeper docs and the few files that define the architecture. Use relative paths.

Write the draft to a temp path so I can review before it replaces the live file: print it to stdout AND save to `./.init-draft.md`.

# PHASE 3 β€” CURATE (I do this, you wait)

Stop after Phase 2. Do not write to <<OUTPUT_FILE>>. Wait for me to review the draft, strip anything a linter or formatter already enforces, replace any inline code snippets with pointers to source files, and only then I will copy the curated content into <<OUTPUT_FILE>>.

# PHASE 3 TAIL FOR Cursor
When I confirm, write the curated content to ./.cursor/rules/project.mdc, with YAML frontmatter `---\ndescription: Project context for Cursor agents\nalwaysApply: true\n---` at the top.
Format quirks: MDC = Markdown with YAML frontmatter. Use `alwaysApply: true` for always-on context, OR `globs: ['**/*.ts']` for path-scoped, OR `description: '...'` + omit alwaysApply for Agent Requested mode.

Cline

.clinerules
File at root, OR directory of *.md files at root (both work; directory enables progressive disclosure). Cline also reads cross-vendor files: .cursor/rules/, .windsurf/rules/, AGENTS.md, CLAUDE.md.
You are bootstrapping a context file for AI coding agents. Do NOT auto-write it. Do this in three phases.

# PHASE 1 β€” SCAN (read-only, no edits)

1. List top-level structure: run `ls -la` then `find . -maxdepth 3 -type d -not -path '*/.*' | sort`.
2. Detect stack: read whichever of these exist β€” package.json, pyproject.toml, requirements.txt, Cargo.toml, go.mod, Gemfile, pom.xml, build.gradle, deno.json, *.csproj.
3. Read existing docs: README*, CONTRIBUTING*, ARCHITECTURE*, docs/index*, .editorconfig.
4. Detect vendor instruction files and per-vendor config already present: CLAUDE.md, AGENTS.md, GEMINI.md, .cursor/rules/*.mdc, .github/copilot-instructions.md, .windsurf/rules/*.md, .clinerules (or .clinerules/ directory), opencode.json.
5. Identify the most-edited files: `git log --since='90 days ago' --pretty=format: --name-only | sort | uniq -c | sort -rn | head -25`.
6. Run `--help` (or read help docs) for any non-obvious tool you find in scripts/Makefile (e.g. `make help`, `bun run`, custom CLI). Do NOT execute build/test commands.

# PHASE 2 β€” DRAFT (write a draft, NOT the final file)

Produce a draft with EXACTLY these sections, in this order. Keep total under 300 lines; target 60.

**WHAT** β€” Stack, languages, frameworks, package manager, runtime versions. One line per item. Only versions you actually verified.

**WHY** β€” Two sentences: what this project is for, and the one non-obvious thing about how it is shaped (monorepo? plugin host? CLI + lib? data pipeline?).

**HOW** β€” Bullet list of exact commands. Install, lint, typecheck, test (unit + integration if separate), build, run dev. Use the actual package manager (bun vs pnpm vs npm vs uv vs poetry vs cargo), not a generic placeholder.

**NON-OBVIOUS** β€” Anything an experienced engineer would not guess from reading the code: required env vars and where to get them, services that must run locally, secrets handling, custom DSLs, generated code that should never be edited by hand, code paths that look dead but are entry points.

**BOUNDARIES** β€” Three tiers in this exact format:
  - βœ… Always OK to modify: <paths/patterns>
  - ⚠️ Ask first: <paths/patterns>
  - 🚫 Never touch: <paths/patterns>

**KEY FILES** β€” Pointers (not copies) to deeper docs and the few files that define the architecture. Use relative paths.

Write the draft to a temp path so I can review before it replaces the live file: print it to stdout AND save to `./.init-draft.md`.

# PHASE 3 β€” CURATE (I do this, you wait)

Stop after Phase 2. Do not write to <<OUTPUT_FILE>>. Wait for me to review the draft, strip anything a linter or formatter already enforces, replace any inline code snippets with pointers to source files, and only then I will copy the curated content into <<OUTPUT_FILE>>.

# PHASE 3 TAIL FOR Cline
When I confirm, write the curated content to ./.clinerules at the repository root.
Format quirks: Plain Markdown. Optional YAML frontmatter with `paths:` field for glob scoping. The tail below writes a single .clinerules file (most common starting point). For larger projects, prefer .clinerules/ as a directory of *.md files β€” Cline reads all of them and you get progressive disclosure (split building/testing/conventions into separate files). If you also want Memory Bank methodology, create the 6 canonical files in `memory-bank/` (projectbrief.md, productContext.md, activeContext.md, systemPatterns.md, techContext.md, progress.md).

OpenAI Codex

AGENTS.md
Walked hierarchically: ~/.codex/AGENTS.md β†’ git root AGENTS.md β†’ subdirectory AGENTS.md. The nearest file to the working code wins. For Codex Cloud also configure your setup script in the cloud environment (runs with internet pre-sandbox).
You are bootstrapping a context file for AI coding agents. Do NOT auto-write it. Do this in three phases.

# PHASE 1 β€” SCAN (read-only, no edits)

1. List top-level structure: run `ls -la` then `find . -maxdepth 3 -type d -not -path '*/.*' | sort`.
2. Detect stack: read whichever of these exist β€” package.json, pyproject.toml, requirements.txt, Cargo.toml, go.mod, Gemfile, pom.xml, build.gradle, deno.json, *.csproj.
3. Read existing docs: README*, CONTRIBUTING*, ARCHITECTURE*, docs/index*, .editorconfig.
4. Detect vendor instruction files and per-vendor config already present: CLAUDE.md, AGENTS.md, GEMINI.md, .cursor/rules/*.mdc, .github/copilot-instructions.md, .windsurf/rules/*.md, .clinerules (or .clinerules/ directory), opencode.json.
5. Identify the most-edited files: `git log --since='90 days ago' --pretty=format: --name-only | sort | uniq -c | sort -rn | head -25`.
6. Run `--help` (or read help docs) for any non-obvious tool you find in scripts/Makefile (e.g. `make help`, `bun run`, custom CLI). Do NOT execute build/test commands.

# PHASE 2 β€” DRAFT (write a draft, NOT the final file)

Produce a draft with EXACTLY these sections, in this order. Keep total under 300 lines; target 60.

**WHAT** β€” Stack, languages, frameworks, package manager, runtime versions. One line per item. Only versions you actually verified.

**WHY** β€” Two sentences: what this project is for, and the one non-obvious thing about how it is shaped (monorepo? plugin host? CLI + lib? data pipeline?).

**HOW** β€” Bullet list of exact commands. Install, lint, typecheck, test (unit + integration if separate), build, run dev. Use the actual package manager (bun vs pnpm vs npm vs uv vs poetry vs cargo), not a generic placeholder.

**NON-OBVIOUS** β€” Anything an experienced engineer would not guess from reading the code: required env vars and where to get them, services that must run locally, secrets handling, custom DSLs, generated code that should never be edited by hand, code paths that look dead but are entry points.

**BOUNDARIES** β€” Three tiers in this exact format:
  - βœ… Always OK to modify: <paths/patterns>
  - ⚠️ Ask first: <paths/patterns>
  - 🚫 Never touch: <paths/patterns>

**KEY FILES** β€” Pointers (not copies) to deeper docs and the few files that define the architecture. Use relative paths.

Write the draft to a temp path so I can review before it replaces the live file: print it to stdout AND save to `./.init-draft.md`.

# PHASE 3 β€” CURATE (I do this, you wait)

Stop after Phase 2. Do not write to <<OUTPUT_FILE>>. Wait for me to review the draft, strip anything a linter or formatter already enforces, replace any inline code snippets with pointers to source files, and only then I will copy the curated content into <<OUTPUT_FILE>>.

# PHASE 3 TAIL FOR OpenAI Codex
When I confirm, write the curated content to ./AGENTS.md at the repository root.
Format quirks: Plain Markdown for AGENTS.md. Command policies (sandbox/approval) live separately in `.codex/rules/` using Starlark DSL. Don't put policy in AGENTS.md β€” it won't be enforced.

OpenCode

AGENTS.md
Multi-source precedence: project AGENTS.md β†’ global ~/.config/opencode/AGENTS.md β†’ opencode.json custom path β†’ CLAUDE.md fallback (disable with OPENCODE_DISABLE_CLAUDE_MD=1). Hierarchical: nearest file to the code wins.
You are bootstrapping a context file for AI coding agents. Do NOT auto-write it. Do this in three phases.

# PHASE 1 β€” SCAN (read-only, no edits)

1. List top-level structure: run `ls -la` then `find . -maxdepth 3 -type d -not -path '*/.*' | sort`.
2. Detect stack: read whichever of these exist β€” package.json, pyproject.toml, requirements.txt, Cargo.toml, go.mod, Gemfile, pom.xml, build.gradle, deno.json, *.csproj.
3. Read existing docs: README*, CONTRIBUTING*, ARCHITECTURE*, docs/index*, .editorconfig.
4. Detect vendor instruction files and per-vendor config already present: CLAUDE.md, AGENTS.md, GEMINI.md, .cursor/rules/*.mdc, .github/copilot-instructions.md, .windsurf/rules/*.md, .clinerules (or .clinerules/ directory), opencode.json.
5. Identify the most-edited files: `git log --since='90 days ago' --pretty=format: --name-only | sort | uniq -c | sort -rn | head -25`.
6. Run `--help` (or read help docs) for any non-obvious tool you find in scripts/Makefile (e.g. `make help`, `bun run`, custom CLI). Do NOT execute build/test commands.

# PHASE 2 β€” DRAFT (write a draft, NOT the final file)

Produce a draft with EXACTLY these sections, in this order. Keep total under 300 lines; target 60.

**WHAT** β€” Stack, languages, frameworks, package manager, runtime versions. One line per item. Only versions you actually verified.

**WHY** β€” Two sentences: what this project is for, and the one non-obvious thing about how it is shaped (monorepo? plugin host? CLI + lib? data pipeline?).

**HOW** β€” Bullet list of exact commands. Install, lint, typecheck, test (unit + integration if separate), build, run dev. Use the actual package manager (bun vs pnpm vs npm vs uv vs poetry vs cargo), not a generic placeholder.

**NON-OBVIOUS** β€” Anything an experienced engineer would not guess from reading the code: required env vars and where to get them, services that must run locally, secrets handling, custom DSLs, generated code that should never be edited by hand, code paths that look dead but are entry points.

**BOUNDARIES** β€” Three tiers in this exact format:
  - βœ… Always OK to modify: <paths/patterns>
  - ⚠️ Ask first: <paths/patterns>
  - 🚫 Never touch: <paths/patterns>

**KEY FILES** β€” Pointers (not copies) to deeper docs and the few files that define the architecture. Use relative paths.

Write the draft to a temp path so I can review before it replaces the live file: print it to stdout AND save to `./.init-draft.md`.

# PHASE 3 β€” CURATE (I do this, you wait)

Stop after Phase 2. Do not write to <<OUTPUT_FILE>>. Wait for me to review the draft, strip anything a linter or formatter already enforces, replace any inline code snippets with pointers to source files, and only then I will copy the curated content into <<OUTPUT_FILE>>.

# PHASE 3 TAIL FOR OpenCode
When I confirm, write the curated content to ./AGENTS.md at the repository root.
Format quirks: Plain Markdown for AGENTS.md. Configuration (custom paths, hooks, model config) goes in opencode.json (JSONC). Hooks are TypeScript plugins.

Antigravity 2.0

GEMINI.md
3-level hierarchy: Global ~/.gemini/GEMINI.md β†’ Workspace GEMINI.md β†’ subdirectories. Antigravity also reads AGENTS.md since v1.20.3 (Mar 2026), but GEMINI.md wins on conflict. Antigravity 2.0 (2026-05-19) supports multi-folder project context β€” one GEMINI.md can govern several roots.
You are bootstrapping a context file for AI coding agents. Do NOT auto-write it. Do this in three phases.

# PHASE 1 β€” SCAN (read-only, no edits)

1. List top-level structure: run `ls -la` then `find . -maxdepth 3 -type d -not -path '*/.*' | sort`.
2. Detect stack: read whichever of these exist β€” package.json, pyproject.toml, requirements.txt, Cargo.toml, go.mod, Gemfile, pom.xml, build.gradle, deno.json, *.csproj.
3. Read existing docs: README*, CONTRIBUTING*, ARCHITECTURE*, docs/index*, .editorconfig.
4. Detect vendor instruction files and per-vendor config already present: CLAUDE.md, AGENTS.md, GEMINI.md, .cursor/rules/*.mdc, .github/copilot-instructions.md, .windsurf/rules/*.md, .clinerules (or .clinerules/ directory), opencode.json.
5. Identify the most-edited files: `git log --since='90 days ago' --pretty=format: --name-only | sort | uniq -c | sort -rn | head -25`.
6. Run `--help` (or read help docs) for any non-obvious tool you find in scripts/Makefile (e.g. `make help`, `bun run`, custom CLI). Do NOT execute build/test commands.

# PHASE 2 β€” DRAFT (write a draft, NOT the final file)

Produce a draft with EXACTLY these sections, in this order. Keep total under 300 lines; target 60.

**WHAT** β€” Stack, languages, frameworks, package manager, runtime versions. One line per item. Only versions you actually verified.

**WHY** β€” Two sentences: what this project is for, and the one non-obvious thing about how it is shaped (monorepo? plugin host? CLI + lib? data pipeline?).

**HOW** β€” Bullet list of exact commands. Install, lint, typecheck, test (unit + integration if separate), build, run dev. Use the actual package manager (bun vs pnpm vs npm vs uv vs poetry vs cargo), not a generic placeholder.

**NON-OBVIOUS** β€” Anything an experienced engineer would not guess from reading the code: required env vars and where to get them, services that must run locally, secrets handling, custom DSLs, generated code that should never be edited by hand, code paths that look dead but are entry points.

**BOUNDARIES** β€” Three tiers in this exact format:
  - βœ… Always OK to modify: <paths/patterns>
  - ⚠️ Ask first: <paths/patterns>
  - 🚫 Never touch: <paths/patterns>

**KEY FILES** β€” Pointers (not copies) to deeper docs and the few files that define the architecture. Use relative paths.

Write the draft to a temp path so I can review before it replaces the live file: print it to stdout AND save to `./.init-draft.md`.

# PHASE 3 β€” CURATE (I do this, you wait)

Stop after Phase 2. Do not write to <<OUTPUT_FILE>>. Wait for me to review the draft, strip anything a linter or formatter already enforces, replace any inline code snippets with pointers to source files, and only then I will copy the curated content into <<OUTPUT_FILE>>.

# PHASE 3 TAIL FOR Antigravity 2.0
When I confirm, write the curated content to ./GEMINI.md at the repository root.
Format quirks: Plain Markdown for GEMINI.md. Workflows use YAML frontmatter with `description`; current Knowledge Subagent behavior is not asserted pending an official locator.

Kiro

AGENTS.md
Kiro's native instruction system is steering files in .kiro/steering/ (foundation trio product.md / tech.md / structure.md, bootstrapped via the Steering panel's 'Generate Steering Docs'). Kiro also reads AGENTS.md at the workspace root automatically, so the universal prompt below targets AGENTS.md for portability; split the content into .kiro/steering/*.md to use Kiro's 4 inclusion modes (always / fileMatch / manual / auto).
You are bootstrapping a context file for AI coding agents. Do NOT auto-write it. Do this in three phases.

# PHASE 1 β€” SCAN (read-only, no edits)

1. List top-level structure: run `ls -la` then `find . -maxdepth 3 -type d -not -path '*/.*' | sort`.
2. Detect stack: read whichever of these exist β€” package.json, pyproject.toml, requirements.txt, Cargo.toml, go.mod, Gemfile, pom.xml, build.gradle, deno.json, *.csproj.
3. Read existing docs: README*, CONTRIBUTING*, ARCHITECTURE*, docs/index*, .editorconfig.
4. Detect vendor instruction files and per-vendor config already present: CLAUDE.md, AGENTS.md, GEMINI.md, .cursor/rules/*.mdc, .github/copilot-instructions.md, .windsurf/rules/*.md, .clinerules (or .clinerules/ directory), opencode.json.
5. Identify the most-edited files: `git log --since='90 days ago' --pretty=format: --name-only | sort | uniq -c | sort -rn | head -25`.
6. Run `--help` (or read help docs) for any non-obvious tool you find in scripts/Makefile (e.g. `make help`, `bun run`, custom CLI). Do NOT execute build/test commands.

# PHASE 2 β€” DRAFT (write a draft, NOT the final file)

Produce a draft with EXACTLY these sections, in this order. Keep total under 300 lines; target 60.

**WHAT** β€” Stack, languages, frameworks, package manager, runtime versions. One line per item. Only versions you actually verified.

**WHY** β€” Two sentences: what this project is for, and the one non-obvious thing about how it is shaped (monorepo? plugin host? CLI + lib? data pipeline?).

**HOW** β€” Bullet list of exact commands. Install, lint, typecheck, test (unit + integration if separate), build, run dev. Use the actual package manager (bun vs pnpm vs npm vs uv vs poetry vs cargo), not a generic placeholder.

**NON-OBVIOUS** β€” Anything an experienced engineer would not guess from reading the code: required env vars and where to get them, services that must run locally, secrets handling, custom DSLs, generated code that should never be edited by hand, code paths that look dead but are entry points.

**BOUNDARIES** β€” Three tiers in this exact format:
  - βœ… Always OK to modify: <paths/patterns>
  - ⚠️ Ask first: <paths/patterns>
  - 🚫 Never touch: <paths/patterns>

**KEY FILES** β€” Pointers (not copies) to deeper docs and the few files that define the architecture. Use relative paths.

Write the draft to a temp path so I can review before it replaces the live file: print it to stdout AND save to `./.init-draft.md`.

# PHASE 3 β€” CURATE (I do this, you wait)

Stop after Phase 2. Do not write to <<OUTPUT_FILE>>. Wait for me to review the draft, strip anything a linter or formatter already enforces, replace any inline code snippets with pointers to source files, and only then I will copy the curated content into <<OUTPUT_FILE>>.

# PHASE 3 TAIL FOR Kiro
When I confirm, write the curated content to ./AGENTS.md at the repository root (or split into ./.kiro/steering/product.md, tech.md, structure.md for Kiro-native steering).
Format quirks: Plain Markdown. Steering files support 4 inclusion modes; foundation files are always loaded. Skills live in .kiro/skills/<name>/SKILL.md (open Agent Skills standard). MCP servers configured in .kiro/settings/mcp.json. No .kiroignore is documented.

Pitfalls β€” what NOT to put in the file

Don't let the agent auto-write the file

/init-style commands write directly. The evaluation paper measures auto-generated files at ~βˆ’3 pp vs no file, and human-curated at ~+4 pp. The 7 pp gap comes from the agent dumping every command it finds; models then start ignoring the bloated file. Phase 3 above keeps you in the loop.

Don't put style rules in the context file

Naming, indentation, formatting β€” those belong in linters/formatters. Putting them in CLAUDE.md / AGENTS.md wastes the ~150-instruction budget and the model often disobeys them anyway. ESLint/Prettier/ruff don't negotiate.

Don't be task-specific

Everything in the file should apply to every session. Task-specific guidance belongs in the prompt you write per task, not in a globally-loaded file.

Prefer pointers to copies

If you need to explain the test layout in depth, write `./agent_docs/testing.md` and reference it. Embedded code snippets in the root context file rot fast and lie to the agent.

Research Pending

Sections scheduled but not yet sourced

These topics belong in the report but currently lack provenance. They are listed here so readers know what is missing rather than guessing. Source Scout will populate them as data becomes available; until then, no fabricated values appear in the matrices.
Performance benchmarks SWE-bench Verified, Terminal-bench, and Aider Polyglot scores per platform β€” including model version and date of run. Needs: Public leaderboard URL + model version + run date per platform
Telemetry & privacy defaults What leaves the user's machine by default. Training-on-data opt-out vs opt-in. Retention window. Citation: official privacy/trust page per vendor. Needs: Privacy/trust page URL + verified statement on default behavior
Compliance certifications SOC 2 Type II, ISO 27001, HIPAA BAA availability, FedRAMP authorization, GDPR/EU residency options. Only Copilot's FedRAMP status is currently in the report. Needs: Trust portal URL + attestation date per cert per platform
MCP server / integration breadth Count of MCP servers in each platform's marketplace or registry (or note 'BYO, no marketplace'). One number is cited for Cline today β€” needs parity for the other seven platforms. Needs: Marketplace URL + count-as-of date per platform
Model coverage matrix Which model families each platform supports today: Claude 4.x, GPT-5.x, Gemini, local via Ollama/vLLM, Mistral, DeepSeek. Scattered across cards now; needs a single matrix. Needs: Vendor docs URL confirming each model family per platform
IDE / OS / language support Matrix of supported IDEs (VS Code, JetBrains, Neovim, web, CLI) Γ— OS (macOS, Linux, Windows). Implied today but never tabulated. Needs: Vendor system-requirements page per platform
Public roadmap Announced-but-not-shipped features per platform, separate from 'What's New'. Currently invisible. Needs: Public roadmap or quarterly note URL per platform
Failure modes Where each agent typically breaks in practice: context-window exhaustion, long-session degradation, cross-file refactor accuracy, tool-call loops. Operational, not marketing. Needs: Mix of public incident reports, vendor docs on limits, and user-community sentiment
Migration & lock-in Cost of switching: do SKILL.md, agent definitions, hooks, and instruction files port across? Where do they not? Needs: Cross-vendor compatibility statements and format specs
New entrant: Grok Build (xAI / SpaceXAI) xAI's Grok Build coding agent + Grok 4.5 (Jul 8, 2026; $2/$6, trained alongside Cursor, offered in Cursor on all plans). Candidate 9th platform. Needs: Official Grok Build docs URL + feature/pricing/model coverage to run the add-a-platform checklist.
New entrant: JetBrains Junie JetBrains Junie coding agent went GA (June 2026), re-architected on ACP. Candidate platform, esp. for the JetBrains ecosystem. Needs: Junie GA feature set + pricing + model coverage.
ACP (Agent Client Protocol) as a cross-vendor standard Zed-created, JetBrains-co-maintained 'LSP for coding agents'; MS Intelligent Terminal 0.1 (Jun 2) shipped a native ACP pane; Junie GA rebuilt on ACP. Cross-platform interop dimension the matrix does not yet capture. Needs: Per-platform ACP support status (client/server) with source URLs.
MCP 2026-07-28 spec revision Biggest MCP revision since launch (RC locked May 21): stateless HTTP core, MCP Apps (server-rendered UI), Tasks (long-running), mandatory OAuth 2.1, new routing headers. SC Media flags new client-state-hijack risks. Needs: Per-platform MCP spec-version adoption once the spec finalizes (2026-07-28).

What's New (vendor changelogs β€” not the report's last-updated date; that is at the top of the page)

Cycle delta

8platforms with updates
76applied
19under review / new
60deferred
9security-related
5pricing-related

Antigravity

vendor: 2026-03-25["How is Gemini changing Maps?", "What is \"vibe design?\"", "How can I learn nenew
vendor: 2026-04-09Try notebooks in Gemini to easily keep track of projectsnew
vendor: 2026-05-19Antigravity 2.0 desktop app launches at Google I/O 2026applied
vendor: 2026-05-19Antigravity CLI launches as new product surfaceapplied
vendor: 2026-05-19Antigravity SDK launches with programmatic agent accessapplied
vendor: 2026-05-19Antigravity integrates with Gemini Enterprise Agent Platformapplied
vendor: 2026-05-19Gemini 3.5 Flash is now the primary model in Antigravity 2.0applied
vendor: 2026-05-19Google AI Ultra $100/mo plan introduced for Antigravityapplied
vendor: 2026-05-19Antigravity 2.0 release-quality concerns: community rollback to 1.23.2applied
vendor: 2026-03-05Antigravity 1.20.3 adds AGENTS.md support, retires command_support, defaults autdeferred
vendor: 2026-06-22IDE, CLI and Antigravity 2.0, all three facing high response timesdeferred
vendor: 2026-06-29- Google Developers Blogdeferred
vendor: 2026-07-14267 lines (226 loc) Β· 34.2 KBapplied
vendor: 2026-07-142026-07-13T23:07:37Zdeferred
vendor: 2026-07-14Cloud Homepagedeferred
vendor: 2026-07-14Antigravity CLI Stuck on "Antigravity Starter Quota"applied
vendor: 2026-08-23Antigravity memory usagedeferred
vendor: 2026-08-23Refresh official Antigravity surfaces/models and withhold unsupported claimsapplied
vendor: 2026-08-24AUG. 24, 2026deferred
vendor: 2026-08-24Prompt Queue is not workingdeferred
vendor: 2026-09-01602 lines (518 loc) Β· 84.9 KBdeferred
vendor: 2026-09-01Black / unreadable text on extension marketplace detail pages when using dark thdeferred
vendor: 2026-09-02612 lines (526 loc) Β· 86 KBdeferred
vendor: 2026-09-02September 2, 2026deferred
vendor: 2026-09-03Cloud Homepagedeferred
vendor: 2026-09-03625 lines (537 loc) Β· 87.2 KBapplied
vendor: 2026-09-03[BUG] All Gemini models fail instantly in Antigravity with HTTP 404 NOT_FOUND, wdeferred
vendor: 2026-09-03R7.1: antigravity pluginsapplied
vendor: 2026-09-03R7.1: antigravity hooksapplied
vendor: 2026-09-03R7.1: antigravity custom agentsapplied
vendor: 2026-09-03R7.1: antigravity worktree isolationapplied
vendor: 2026-09-03R7.1: antigravity remote controlapplied

OpenAI Codex

vendor: 2026-03-15Clarify Codex agent-definition wording so AGENTS.md is not mistaken for an agentapplied
vendor: 2026-03-180.116.0-alpha.6deferred
vendor: 2026-03-18Docs, videos, and demo apps for building with OpenAIdeferred
vendor: 2026-03-19tagorr, breakcraft, yougrandpa, and potapenko reacted with rocket emojideferred
vendor: 2026-03-250.117.0-alpha.17new
vendor: 2026-03-25Notebook examples for building with OpenAI modelsnew
vendor: 2026-04-09Security and qualitynew
vendor: 2026-04-09Guides, concepts, and product docs for Codexnew
vendor: 2026-06-29Release notes from codexdeferred
vendor: 2026-06-29Workspace Agentsdeferred
vendor: 2026-07-142026-07-14T07:11:23Zdeferred
vendor: 2026-07-14Example workflows and tasks teams can take on with ChatGPT or Codexapplied
vendor: 2026-08-23Refresh Codex custom agents, skills, and stable releaseapplied
vendor: 2026-09-01Mutual TLS (mTLS)deferred
vendor: 2026-09-02ChatGPT Workdeferred
vendor: 2026-09-03September, 2026deferred
vendor: 2026-09-03R7.1: codex pluginsapplied
vendor: 2026-09-03R7.1: codex local memoryapplied

Kiro

vendor: 2026-07-142026-07-14T04:29:29.219Zapplied
vendor: 2026-08-23Refresh Kiro hooks, advisories, subagents, models, and releasesapplied
vendor: 2026-09-01System & storagedeferred
vendor: 2026-09-02Configuration Syncdeferred

GitHub Copilot

vendor: 2026-03-15Clarify Copilot handoff scope across IDE and GitHub.com surfacesapplied
vendor: 2026-03-18Improvementapplied
vendor: 2026-03-25Improvementnew
vendor: 2026-04-09Copilot-reviewed pull request merge metrics now in the usage metrics APInew
vendor: 2026-06-29https://github.blog/changelog/label/copilot/applied
vendor: 2026-07-14Tue, 14 Jul 2026 13:02:40 +0000applied
vendor: 2026-08-23Refresh Copilot handoff boundary, Free plan, and security reviewapplied
vendor: 2026-08-24Mon, 24 Aug 2026 17:28:13 +0000deferred

Cursor

vendor: 2026-03-25Mar 19, 2026new
vendor: 2026-04-09Apr 8, 2026new
vendor: 2026-06-29What's New in Cursor β€” Latest Updates & Release Notesapplied
vendor: 2026-07-14Automationsapplied
vendor: 2026-08-23Refresh current Cursor platform delta and withhold ungrounded CVEunder review
vendor: 2026-09-01For Origin-hosted repos, Origin is the source of truth. Pushes land on Origin, aapplied
vendor: 2026-09-03Sep 2, 2026applied
vendor: 2026-09-03R7.1: cursor pluginsapplied
vendor: 2026-09-03R7.1: cursor agents windowapplied
vendor: 2026-09-03R7.1: cursor remote mobileapplied
vendor: 2026-09-03R7.1: cursor security agentsapplied
vendor: 2026-09-03R7.1: cursor setup scriptapplied
vendor: 2026-09-03R7.1: cursor memory withheldapplied

Claude Code

vendor: 2026-03-15Operational example: review stale voice-input rollout wording before future releunder review
vendor: 2026-03-18March 17, 2026applied
vendor: 2026-03-25Use Claude Codenew
vendor: 2026-04-09Explore the .claude directorynew
vendor: 2026-06-29Claude Code changelog - Claude Code Docsapplied
vendor: 2026-07-14Chrome extensionapplied
vendor: 2026-08-23Refresh Claude Code voice, hooks, and access surfacesapplied
vendor: 2026-09-01August 31, 2026deferred
vendor: 2026-09-02September 1, 2026deferred
vendor: 2026-09-03September 2, 2026applied
vendor: 2026-09-03R7.1: claude local code reviewapplied
vendor: 2026-09-03R7.1: claude auto memoryapplied

Cline

vendor: 2026-03-2520 Mar 23:27new
vendor: 2026-04-09Security and qualitynew
vendor: 2026-06-29GitHub Copilot appapplied
vendor: 2026-07-14Release listapplied
vendor: 2026-08-23Choose a tag to comparedeferred
vendor: 2026-08-23Refresh Cline pricing, repository, plugins, and subagentsapplied
vendor: 2026-09-01Desktop v0.0.21deferred
vendor: 2026-09-02SDK v0.0.82deferred
vendor: 2026-09-03Desktop v0.0.23-beta.1applied
vendor: 2026-09-03R7.1: cline pluginsapplied

OpenCode

vendor: 2026-04-09Security and qualitynew
vendor: 2026-06-29Release notes from opencodeapplied
vendor: 2026-07-142026-07-13T21:09:41Zapplied
vendor: 2026-08-23Correct OpenCode repository, release, hooks, and CVE factsapplied
vendor: 2026-09-03R7.1: opencode isolation withheldapplied

Glossary

Agent
An LLM-driven loop that can call tools, edit files, and run commands on the user's behalf. The catch-all label every vendor uses; specifics differ wildly.
compare: subagent, skill
Subagent
An agent spawned by another agent with a narrower scope or different model. Claude Code's named markdown definitions are the reference implementation; Copilot calls these "handoffs", Cursor calls them "/multitask".
see also: Agent Teams, peer-to-peer messaging
Skill (SKILL.md)
A reusable, file-based capability following the SKILL.md open standard: triggers, instructions, optional code. Originated in Claude Code; partial support spreading to other platforms.
see also: instruction file
Hook
A user-defined script or rule triggered at lifecycle points (pre-tool, post-tool, on-commit, on-stop). Lets you enforce policy, run validators, or shape behavior deterministically.
compare: lifecycle event, policy
MCP (Model Context Protocol)
Anthropic-originated open standard for connecting LLMs to external tools, data sources, and APIs over a uniform interface. Adopted by most platforms with varying server registries.
see also: tool, integration
Lens
A filter view in this report (Security / Cost / Open Source) that dims content not relevant to that concern. Not a vendor feature.
report-internal term
Executive vs Deep view
Report view modes. Executive shows decision-critical sections only; Deep exposes the full matrix, agent catalogs, and configuration tables.
report-internal term
Instruction file
The per-project memory file an agent reads on startup (CLAUDE.md, .cursorrules, .github/copilot-instructions.md, AGENTS.md). Hierarchy and merge semantics vary.
see also: project memory
Auto-delegation
When the primary agent autonomously routes a subtask to a different model, agent, or tool without asking the user. Carries cost and security implications.
see also: permission mode
Permission mode
Sandbox / approval policy controlling which tools the agent can run unattended. Names vary: "plan", "acceptEdits", "dontAsk", "bypassPermissions" in Claude Code; "yolo", "safe", "auto" elsewhere.
see also: hook
BYOK
Bring Your Own Key. The platform runs against a model provider account that you pay directly, instead of bundling inference into a subscription.
see also: pricing model
Confidence (this report)
A 0.0-1.0 score on each tracked claim reflecting source authority, source count, and recency. Surfaced as colored dots and per-section averages.
see also: claim, last verified
Severity (risks)
Categorical impact rating: critical (CVSS 9-10 / RCE / unauthenticated), high (CVE / exfiltration / prompt injection), medium (cost or operational), low. Explicit field per item, with a text-pattern fallback.
see also: prevalence
Prevalence (risks)
How widely the risk is documented: confirmed (CVE/incident with public PoC), observed (multiple reports, no formal disclosure), theoretical (plausible by design, no public report).
see also: severity
Two-layer architecture
This report separates product capability (feature matrix, subagents, config) from market evidence (enterprise adoption cases). Enterprise cases are signals, not normalized feature scores.
report-internal term

Methodology

Data Sources

14 platform changelogs and release pages monitored. Sources fetched daily via automated pipeline. Snapshots stored and diffed for change detection.

Scoring

Dimension scores computed from registry data: βœ… = 1.0, ⚠️ = 0.5, ❌ = 0.0, averaged across all rows in each section. Overall = mean of 5 dimensions (Features, Customization, Skills, Onboarding, Execution).

Confidence

84 claims tracked with confidence scores (0.0–1.0). Confidence based on: number of independent sources, source authority (official docs > community reports), and recency of verification.

Harness Engineering

Harness recommendations use a separate assertion and evidence namespace. They are excluded from the coding-tool claim count and render only after their strict provenance contract passes.

Update Cycle

Daily: automated source fetch. Weekly: change candidate review. Per-release: human adjudication β†’ patch context β†’ adversarial review β†’ patch application. Every change requires human approval.

Limitations

Single-author perspective. English-only sources. No user surveys or telemetry. Snapshot-based (not real-time). Pricing may be outdated within days. Beta/preview features may change without notice. Security disclosures may have embargoed details.
Research, source access and build dates
Data verified through
Not established for the whole edition. Latest recorded active-claim verification: 2026-09-09; oldest: 2026-06-19. These legacy dates do not prove every public assertion was reviewed.
Sources checked at
Not available in the versioned public build. Host-local acquisition runs are separate from claim verification; no successful monitoring run is inferred.
Rendered at (declared artifact input)
2026-09-10T10:15:00Z
Deployed at
Not embedded in this artifact. The exact production receipt is retained separately; this rendered file cannot attest to its own later deployment.
Review deadlines evaluated at

Source acquisition may require manual review even when a claim was substantively reverified. Fetch failures do not disprove a fact, and successful fetches do not reverify it. Per-source acquisition health is unavailable in this build.

Deadlines use the declared evaluation time; opening this static page later can only downgrade displayed review status.