AI Coding Agent Pricing in 2026: Billing Models, Cost Variables, and How to Budget

Published August 6, 2026 · Updated September 2, 2026By ABD Legacy LLC
AI agents / pricing

AI coding agent pricing is not the per-seat number on the pricing page. That is the central claim of a detailed guide TrueFoundry published on August 6, 2026 — and it matches what agencies see when their clients' first coding-agent invoices land. Flat, credit, and pay-per-token plans respond differently to the same usage, and model choice can move the bill more than seat count does.

This page distills the guide into the four billing models — including Grok Bot's new subscription-bundled agent access — the seven cost variables that predict spend, what pricing pages omit, and a six-step budgeting checklist you can hand to finance. It also covers Slack Code, the agentic-coding channel Salesforce launched August 20, 2026 that makes Claude, GitHub Copilot, Devin, and Vercel agents accessible from collaborative Slack channels. If you need to price agency delivery of AI-assisted work rather than internal seats, we cover that too — including what the August 2026 DHH–Lex Fridman debate on vibe coding vs agentic engineering means for projects vs retainers. And if you're choosing between Claude Code and Codex CLI — or deciding which harness your agency should standardize on — the August 2026 agent harness rankings are below.

The four AI coding agent billing models

Every major tool — Cursor, GitHub Copilot, Windsurf, Claude, OpenAI Codex, and now Grok Bot — maps to one of four billing structures, per TrueFoundry's August 6, 2026 guide:

Billing modelExamplesHow it worksWhat to watch
Flat per-seat subscriptionCursor Pro, Windsurf Pro, Claude ProFixed monthly fee per developer; usage limits may exist but the primary cost is predictable.Overage charges when limits are exceeded — and limits often aren't published clearly.
Seat plus creditsGitHub Copilot Pro, Pro+, MaxLower base fee per seat plus a monthly pool of AI credits; stay inside the pool and costs stay contained.Credit burn rates by model — a frontier model can consume credits several times faster than a standard model on the same task.
Pay-per-token APIClaude Code (API mode), OpenAI Codex, Muse Code (beta)No per-seat charge; billing follows token consumption at published API rates.Token volume is harder to estimate than seat count; teams without usage visibility routinely misbudget API-based plans, and sustained high-volume agent workloads can make it the most expensive option. Newest entrant Muse Code (beta, no GA date confirmed) enters at $1.25/M input and $4.25/M output, with a data-sharing contributor tier and enterprise zero-data-retention starting.
Subscription-bundled agent accessGrok Bot (SpaceXAI) — SuperGrok Heavy, Cursor Ultra, Cursor Teams PremiumNo per-token meter and no standalone price: always-on agents run 24/7 on their own cloud computer and are bundled into eligible subscription tiers. Beta, desktop (macOS) and iOS; enterprise waitlist.No standalone pricing announced, so apples-to-apples comparison is impossible until GA; authentication mechanics are undisclosed; beta access is gated by the eligible tiers — and the Cursor deal behind two of them closed August 14, 2026 (SpaceX's $60B all-stock acquisition of Cursor), after which OpenAI said it will wind down OpenAI model supply to Cursor effective November 12, 2026.

Grok Bot also breaks the tool-interop assumption. SpaceXAI says its bots can sign into and work across apps, tools, and websites — including platforms with no clean API or MCP — while Claude Code, OpenAI Codex, and Muse Code rely on MCP (and documented APIs) for tool interop. That matters for the budget conversation: an agent that works where the work lives can replace more manual steps, but with no standalone price and undisclosed authentication mechanics, agencies should treat it as a beta evaluation, not a budget line item.

"Pricing pages show the entry point. Invoices show consumption. Budget for the second, not the first." — TrueFoundry, AI Coding Agent Pricing (Aug 6, 2026)

Slack Code: AI coding agents, now inside collaborative Slack channels

On August 20, 2026, Salesforce launched Slack Code — dedicated "code channels" inside Slack where teams and AI coding agents build software together. Slack frames it as "multiplayer" agentic coding: mention an agent like Claude or Devin in a conversation, and it spins up a project-specific code channel, pulling in relevant teammates and context. The channel shows code diffs, live HTML previews, and planning docs; anyone in the channel can review, give feedback, pause, redirect, or stop the agent, and approve work before it ships. When the work is done, the channel archives itself while the history stays searchable, like an audit log.

Availability: Slack Code is available on any Slack plan from day one, with four agent integrations live at launch: Claude (Anthropic), Devin (Cognition), GitHub Copilot, and Vercel. OpenAI's ChatGPT is a founding partner but was "available soon" at launch — the only named agent not yet usable in code channels. Slack says more partners are on the way, and it is pitching code channels beyond engineering: marketing campaigns, legal document review, IT onboarding.

Pricing impact: Slack Code itself has no separate price — it ships inside Slack plans. But per Salesforce, "access to each partner agent is required": agencies still pay for Claude, Copilot, Devin, or Vercel subscriptions on top of Slack seats. It is a distribution layer, not a new billing model — the four billing structures above stay valid; what changes is where agent work happens and who sees it.

Why it matters for agency decision-making:

Analysis: the implications above are editorial inference from the sourced launch facts, not claims made by the sources. Slack Code does not change what coding agents cost; it changes where agent work happens — and who sees it. Teams already on Slack gain a zero-marginal-cost surface for agent delivery; the marginal cost remains the per-agent subscription, which the pricing math above still governs.

See our Slack Code launch brief for full coverage of what agencies should do next.

Claude Code weekly limits change September 14: +25% permanent baseline, net −17% vs today

On August 29, 2026, Anthropic announced that starting September 14, 2026 it is permanently raising Claude Code's standard weekly usage limits by 25% over the pre-promotion baseline for Pro, Max, Team, and seat-based Enterprise plans — and the temporary 50% weekly boost (active since May 13, 2026, extended through the summer) expires the same day (@ClaudeDevs, Aug 29, 2026; BleepingComputer). Anthropic confirmed the net effect itself: "Compared to today, this works out to a 17% reduction in weekly limits on Claude Code."

PeriodClaude Code weekly limit (relative to pre-promotion baseline)
Pre-promotion baseline100%
Today — temporary +50% boost (through September 13, 2026)150%
From September 14, 2026 (permanent)125%

Going from 150 to 125 is a 16.7% cut (25 ÷ 150), which Anthropic and every major outlet round to about 17%. A workload consuming 100% of today's boosted cap will consume 120% of the new cap (150 ÷ 125 = 1.2) starting September 14.

Plan scope: the change applies to Pro, Max (5x and 20x), Team (standard and premium seats), and seat-based Enterprise plans. Free and consumption-based (pay-as-you-go) Enterprise seats are excluded. It is a weekly-capacity change only — five-hour session limits, plan pricing, Claude chat, Claude Cowork, and Console API pay-per-token pricing are unchanged. For Pro and Max, Claude apps and Claude Code draw from the same weekly bucket; Team plans give individual member limits, not an org-wide pool.

One documentation note: as of August 30, 2026, Anthropic's Help Center promo article still lists the 50% boost as running through August 31 at 11:59 PM PT; the September 14 date comes from Anthropic's announcement on X. Plan on the September 14 schedule, since it is the company's stated date.

For agencies running Claude agents at scale, three implications:

See Claude Code limits drop 17% on Sept 14 — what agencies should do for the full breakdown and the five pre-change actions.

The seven cost variables that predict spend

TrueFoundry's guide argues six factors predict spend better than the headline number — and OpenAI's August 11, 2026 agent-import wave added a seventh: migration and import effort. Together:

#VariableWhy it matters
1Usage volumeLight autocomplete vs. all-day agent workflows can differ 20–50× in token consumption.
2Model selectionOn credit plans, premium models can burn credits up to 8× faster than standard ones.
3Billing modelFlat, credit, or API — the same usage profile produces three different cost outcomes.
4Team sizeWhat's economical at 10 developers isn't always economical at 40.
5Billing cycleAnnual discounts (often 15–20%) change total cost of ownership.
6Plan tier fitFree and entry tiers often carry quotas production teams exceed within weeks.
7Migration / import effortSince OpenAI's Aug 11, 2026 import wave, moving an established Claude Code or Cursor setup into ChatGPT Work/Codex is an import plus a review pass instead of a days-long rebuild — but post-import validation (permissions, MCP auth, hooks, placeholder templates, plugin re-auth) is still billable work, and import is one-way into OpenAI.

What pricing pages tend to omit

The working example in the source is stark: a task that costs one credit unit on a standard model might cost eight on a frontier model. Budget against that spread, not against the cheapest option on the pricing page.

The coding-agent price war is starting to bend both spread variables: Meta launched Muse Code (beta, no GA date confirmed) on Aug 5, 2026 at $1.25/M input and $4.25/M output — undercutting Claude Code API and OpenAI Codex at volume — so the 20–50× usage spread and the 8× credit-burn rate become bigger levers in the budget conversation when teams switch models to chase cheaper tokens.

Budgeting for AI coding agents: a six-step checklist

  1. Map your team's usage profile. Headcount first, then intensity: light (autocomplete, occasional prompts), moderate (daily agent-assisted tasks), or heavy (continuous agent workflows across large codebases). Model a range if uncertain — the gap between low and high estimates is your budget risk exposure.
  2. Match the billing model to your priorities. Stable monthly number → flat per-seat (you may overpay for light users). Lower entry plus flexibility → seat-plus-credits (needs model governance or overages surprise you). Pay-only-for-use → pay-per-token API (harder to forecast; spikes happen). There is no universal right answer.
  3. Sanity-check free tiers before anyone gets attached. Compare the quota to expected per-developer usage, confirm per-user vs. shared quota, and find out what happens when exceeded — hard stop, throttling, or silent overage.
  4. Set model governance before rollout. Define defaults for routine work, approve premium models only for genuinely hard problems (with criteria), and review monthly during adoption, quarterly once things stabilize. Model selection is a budget decision disguised as a settings preference.
  5. Plan for scale explicitly. The cheapest option for a 10-person pilot is often not the cheapest at 40 developers. Re-run the evaluation when team size or usage shifts.
  6. Align with finance before you commit. Provide projected cost under low, moderate, and high usage scenarios, a plain-language explanation of the billing model and its variability drivers, and a spend review plan — TrueFoundry's teams use a 90-day checkpoint.

The transparency case: why the invoice doesn't match the plan

TrueFoundry's motivating example is the one agencies hear constantly: a team buys a "$20 per seat per month" plan, and six weeks later someone asks why the invoice doesn't match the plan everyone thought they bought. Finance budgets from per-seat estimates; AI coding agent invoices reflect consumption. That mismatch causes approval and reconciliation problems downstream.

Transparency changes behavior: teams that see usage and spend data change how they use agents more than any policy doc does. The practical mitigations are the same as for agency delivery cost — budget the loops, not just the tokens, govern model selection, and insist on budget rails before rollout.

Agent harness rankings (August 2026): Claude Code vs Codex CLI

Claude Code ranks #1 and Codex CLI ranks #2 in the August 2026 agent harness rankings — and the deciding layer is the harness, not the model. Machine Brief's rankings (August 24, 2026) put Claude Code first, Codex CLI second, Cursor third, with Gemini CLI and GitHub Copilot rounding out the top five. explainx.ai's updated top-10 list (published July 22, 2026; updated August 24, 2026) agrees on the top two and fills out the field: Claude Code, Codex CLI/ChatGPT Work, Cursor, Google Antigravity, GitHub Copilot (agent mode), Devin, Windsurf, Factory (Droid), Amazon Q Developer, and Replit Agent.

RankHarnessMachine Brief (Aug 24, 2026)explainx.ai (updated Aug 24, 2026)
1Claude CodeFirst on depth of hooks, subagents, and dynamic workflows; default choice for long autonomous coding sessions.#1 closed-source — the reference implementation most other harnesses are compared against.
2Codex CLILeads on cloud, pull-request-shaped autonomy.#2 closed-source (Codex CLI / ChatGPT Work).
3CursorHolds in-editor workflows — but OpenAI model supply ends November 12, 2026 (OpenAI notice Aug 28, 2026); Claude, Gemini, and Grok models remain.#3 closed-source; no access to OpenAI's upcoming Astra model.
4Gemini CLI / AntigravityGemini CLI rounds out the top five.#4 — Google Antigravity, which explainx says replaced Gemini CLI as Google's primary coding-agent surface.
5GitHub CopilotRounds out the top five; unaffected by the Cursor–OpenAI split — Microsoft keeps OpenAI models for Copilot.#5 — GitHub Copilot (agent mode).
"Claude Code ranks first on depth of hooks, subagents and dynamic workflows, and is named the default choice for long autonomous coding sessions. Codex CLI leads on cloud, pull-request-shaped autonomy. Cursor holds in-editor workflows, with Gemini CLI and GitHub Copilot rounding out the top five." — Machine Brief, August 24, 2026

Why the top two matter for budgeting: they run the same frontier models. "Several of these tools run the same frontier models. Claude Code and Codex CLI aren't competing over whose underlying brain is smarter. What differs is session management and how gracefully they fail" (Machine Brief, Aug 24, 2026). The same report cites Linear telemetry showing coding agents tripled pull requests without cutting cycle time — "a better harness raises throughput without raising speed" — because the bottleneck sat in review, not generation. The pricing takeaway is direct: you pay for the harness in throughput and for the model in tokens.

The rankings also admit what they don't measure: "They score depth of hooks, subagents and workflow dynamics. They don't measure the things that actually eat a quarter: onboarding friction, cost per session, or what happens to your repo when an agent goes wrong and nobody notices." Treat the ranking as the shortlist — and the cost-per-session math above as the budget line.

Cursor vs Claude Code vs Copilot: what the OpenAI split changes

On August 28, 2026, OpenAI notified SpaceX that it intends to wind down the contract providing OpenAI models to Cursor, with a proposed shutoff date of November 12, 2026 — after SpaceX completed its $60B all-stock acquisition of Cursor on August 14, 2026. OpenAI also said it will not provide future models to Cursor, including its upcoming Astra model, citing an inability to be confident SpaceX will use its technology within OpenAI's terms of service (OpenAI, Aug 28, 2026; Business Insider, Aug 28, 2026; TNW, Aug 29, 2026).

Cursor says OpenAI models served about 5% of its user traffic. Cursor co-founder and CEO Michael Truell responded on August 29, 2026 that "OpenAI models serve about 5% of Cursor user traffic, and we're speaking with the OpenAI team to resolve this" — meaning the other ~95% already flows through Anthropic Claude, Google Gemini, and xAI Grok, including Grok 4.5, which Cursor and SpaceXAI shipped in July 2026 and which is available on every Cursor plan (wccftech, Aug 29, 2026; Cryptobriefing, Aug 29, 2026; Cursor docs).

Anthropic is expanding Claude compute for Cursor, not retreating. Anthropic co-founder and Chief Compute Officer Tom Brown said "we'll continue to increase compute to support Claude models in Cursor," building on the existing Anthropic–SpaceX Colossus 1 compute partnership announced May 6, 2026 (wccftech, Aug 29, 2026; Anthropic, May 6, 2026). For agency budgeting, the practical read: Claude capacity in Cursor is expanding in the near term even as OpenAI models leave.

FactorCursorClaude CodeGitHub Copilot
OpenAI model access (Aug 2026)Ends November 12, 2026; no Astra access; GPT-5.6 family still listed todayN/A — Anthropic-nativeUnaffected — Microsoft retains OpenAI models
Model mix after the splitClaude, Gemini, Grok (Grok 4.5 on every plan); ~95% of traffic already non-OpenAI per CursorClaude (Opus 5, Sonnet 5, Fable 5.1)OpenAI models via Microsoft; multi-model options
Supply risk from the splitHigh — OpenAI wind-down on Nov 12, 2026Low — Anthropic expanding compute (Colossus 1 + new pledge)Low — Microsoft–OpenAI agreement unchanged
Pricing signal (Aug 2026)Pro $20/mo; Pro+ / Ultra 3x/20x limits; usage-based overagePro $20/mo, Max $100/$200; API per-tokenPro $10/mo; Pro+, Max credit tiers
August 2026 harness rank#3 (Machine Brief; explainx.ai)#1 (Machine Brief; explainx.ai)#5 (Machine Brief; explainx.ai)

What the verdict changes for agencies: Cursor remains a strong in-editor harness — but its OpenAI-model workflows have a hard end date. Agencies should audit which Cursor workflows actually depend on OpenAI models (Cursor's own claim is ~5% of traffic, and per-team mixes vary), test Claude Code and Copilot on their own repos before November 12 rather than switching blind, and re-price per-developer model spend now that Claude compute in Cursor is expanding and OpenAI access is shrinking. See the full timeline in OpenAI Is Cutting Off Cursor: What It Means for AI Coding Agents, and run the per-developer cost math on aiagencycalculator.com/cursor-model-cost-shift.

Claude Code vs Codex CLI: strengths, limitations, and best-use cases

Claude Code is "the reference implementation most other harnesses get compared against" (explainx.ai): terminal-first, with /loop for autonomous retry cycles, subagent delegation, hooks for lifecycle events, and CLAUDE.md as a persistent project-memory file the harness reads on every session. Its permission model — prompting before file writes and shell commands unless explicitly trusted — set the default UX pattern most competitors now copy.

Codex CLI "leads on cloud, pull-request-shaped autonomy" (Machine Brief). It's now folded into a single desktop app alongside Chat and Work modes, sharing usage limits across all three: Work mode leans toward broader "deliverable" tasks — documents, spreadsheets, and reports — while Codex mode stays scoped to repositories (explainx.ai).

Budgeting for a Codex-based workflow? See our Codex CLI pricing guide — Plus, Pro 5x, Pro 20x, Business, and API tiers with usage limits and banked resets.

Which agent harness should agencies standardize on?

Both sources reframe the buying decision the same way: choose the harness layer — context management, subagent coordination, failure handling — before you argue about models. explainx.ai's practical guidance:

Decision factors, per explainx.ai: compliance and data residency, team size, model flexibility, customization depth, and budget shape. One naming caveat before you standardize on Google's entry: Machine Brief (Aug 24) still lists Gemini CLI in its top five, while explainx says Google Antigravity "replaced Gemini CLI as Google's primary coding-agent surface." Both readings are current as of late August 2026 — check which surface Google is actually shipping before you bet an agency workflow on it.

Agentic engineering vs vibe coding: what the DHH–Lex debate means for pricing and engagement models

On August 26, 2026, DHH — creator of Ruby on Rails, CTO of 37signals, and founder of the Omarchy Linux distribution — spent five hours on Lex Fridman Podcast #501 arguing that the industry's real divide is no longer human vs AI coding, but vibe coding vs agent-accelerated engineering. For AI coding agents for agencies, the episode is a pricing brief: vibe coding is the cheap tier, agent-accelerated engineering the premium one. DHH's operational definition: vibe coding is telling an agent to build software and not looking at the implementation; agent-accelerated development keeps the human engaged with what the agent produces (0:51:05). We cover the full argument, its pushback, and the "will AI replace programmers" question in Agentic Engineering vs Vibe Coding: How AI Coding Agents Are Rewriting Agency Delivery. This section is about what the distinction does to your quote.

When vibe coding is acceptable — and what to charge

Vibe coding is legitimate for scoped, low-risk work: prototypes, internal tools, marketing pages, throwaway scripts, and greenfield experiments where architectural decay is acceptable. DHH himself keeps vibe-coded projects. It is cheap to deliver — agents implement and nobody reviews deeply — and it should be priced that way: fixed fee, fast turnaround, disposable output. The failure mode is scale: when "vibe" output accumulates on a production system, architecture decays (his Basecamp 5 story — designers shipping PRs that "destroyed the architecture of the system", 0:16:44) and someone pays for the cleanup, usually more than the vibe saved.

Why agentic engineering commands premium rates

The premium is not the agent — it is the layer around it. DHH's explanation of why the 10X–1000X productivity claims only materialize when humans don't slow the loop:

"As soon as you're having human teams work together on something, the bottleneck is rarely implementation. It's human bandwidth and communication. … to get that magical 10X, 100X, in a few rare cases, 1000X productivity boost, you have to interact with the agents directly, and you cannot intermediate that bandwidth with another human because it's simply too slow." — DHH, Lex Fridman Podcast #501 (0:18:44), transcript

Two pricing implications follow. First, projects priced on "agent hours" are pricing the wrong unit: the billable unit is the review, the architecture call, and the sign-off. Second, DHH's stated operating procedure is two-model cross-review — one frontier model does the work, a differently sourced model reviews it — which adds a second token bill but catches one vendor's agent blind spots. On production work, that second-model review is a real cost line agencies should quote explicitly rather than absorb.

Projects vs retainers: how the debate settles the engagement model

ModeRight workTypical structurePrice signal
Vibe codingPrototypes, internal tools, one-off scripts, greenfield demosFixed-fee project; fast, disposableLow. Priced as output, not as a system.
Agent-accelerated development (agentic engineering)Production systems, existing codebases, anything with a second user or a renewal dateRetainer or sprint-plus-review; continuous direction and auditPremium. Priced on the review and governance layer, plus token cost.

DHH's own workflow is the retainer argument: agents drafted or closed PRs on a regular schedule, and he made "the final determination" on a shortlist (2:43:27). Continuous work needs a continuous engagement — a retainer with a defined review capacity — because the agent work never stops between milestones. A project quote assumes work ends at delivery; agentic engineering doesn't, and the pricing should say so.

Refreshing the "end of manual programming" claim

Earlier framing of this debate treated AI coding as simple substitution — manual programming ends, agents take over, agencies price accordingly. DHH's actual position is more precise, and more useful for pricing. He argues the economic case for hand-sweating every line is eroding — not the value of builders:

"The reason why I, for 25 years, was sweating every line of code so judiciously was because I knew the payoff of keeping an architecture coherent and malleable was software that could change and evolve quickly with a small team … That was premised on humans doing the modifications. I think it is an open question to which degree this still matters." — DHH, Lex Fridman Podcast #501 (1:00:35), transcript

And the romanticization of handwritten code is already underway:

"It's already happening, the romanticization of handwritten code. And I have some of it because it was a very romantic era. I'm grateful to have been alive for 20 years of economically valuable handwritten code. That was a good time." — DHH, Lex Fridman Podcast #501 (1:04:57), transcript

What this means for pricing: don't sell "handwritten" as a premium, and don't discount "agent-written" as a commodity. The premium is engineering judgment — what to build, what to let the agent build, what to review, and how to keep the system coherent while agents move fast. One counterintuitive implication from the episode: over-prescription damages agent output. DHH on his own workflow: "more often than not, I have the humility to recognize that the agent knows best" (0:55:30, transcript) — so the premium lies in problem definition and outcome review, not in dictating the implementation route. Clients shouldn't pay for heavy-handed specs; they should pay for someone who knows which details to specify and which to leave to the agent.

Analysis: DHH never says "agency" or "client" in the episode; the engagement-model mapping above is our extrapolation of his adoption advice. All quotes are verbatim from the official Lex Fridman Podcast #501 transcript (Aug 26, 2026), with YouTube deep links.

Run the numbers for your agency or client work

Use the free AI Agency Pricing Calculator to model setup fees, retainers, and margins — including open-weight vs. frontier model cost profiles.

Open the AI Agency Pricing Calculator →

Frequently asked questions

How much does an AI coding agent cost per month?

Entry-level plans run about $10–$20 per seat per month (GitHub Copilot at $10; Cursor, Claude Code, Windsurf, and OpenAI Codex around $20). The real cost is billing model × usage: light users can stay near the base fee, while all-day agent workflows consume 20–50× more tokens and commonly land in the $100–$400+ range on credit or API plans.

What are the AI coding agent billing models?

Four: flat per-seat subscription (Cursor Pro, Windsurf Pro, Claude Pro), seat plus credits (GitHub Copilot Pro, Pro+, Max), pay-per-token API (Claude Code API mode, OpenAI Codex, Muse Code beta), and subscription-bundled agent access (Grok Bot — no per-token meter, no standalone price, included with SuperGrok Heavy, Cursor Ultra, or Cursor Teams Premium in beta). Each responds differently to the same usage.

Why is my AI coding agent bill higher than the per-seat price?

Pricing pages show the entry point; invoices show consumption. Overages on credit plans, premium-model credit burn (up to 8× faster), token consumption on API plans, and free-tier quotas teams exceed within weeks all push real spend above the sticker price.

How much does model choice change AI agent cost?

A lot. On credit plans, premium models can burn credits up to 8× faster than standard ones — a task that costs one credit unit on a standard model might cost eight on a frontier model. The model picker is a cost lever as well as a settings control.

How do I budget for AI coding agents?

Map the usage profile first, then match the billing model to your priorities, sanity-check free tiers, set model governance before rollout, plan for scale, and give finance a low/moderate/high scenario range with a 90-day spend review checkpoint.

Which AI coding agent billing model is cheapest?

There is no universal answer. Flat per-seat suits teams that want a stable monthly number; seat-plus-credits suits teams that can stay inside a credit pool with model governance; pay-per-token API suits light or uneven usage but can become the most expensive option under sustained high-volume agent workloads; and subscription-bundled agent access (Grok Bot) is the cleanest "no meter" option for teams already on an eligible tier — though its standalone economics aren't priced yet.

How much does Grok Bot cost?

No standalone price has been announced. Grok Bot is in beta and included with eligible subscriptions — SuperGrok Heavy, Cursor Ultra, and Cursor Teams Premium — on desktop (macOS) and iOS, with an enterprise waitlist. SpaceXAI hasn't disclosed how bot authentication works, and the Cursor deal behind two of the eligible tiers closed August 14, 2026: SpaceX completed the $60B all-stock acquisition of Cursor, and OpenAI has since said it will wind down OpenAI model supply to Cursor effective November 12, 2026.

How much does Slack Code cost?

Slack Code has no separate price — it launched August 20, 2026 on any Slack plan. But "access to each partner agent is required": Claude, GitHub Copilot, Devin, and Vercel agents are billed through their own subscriptions on top of Slack seats, and OpenAI's ChatGPT is "available soon." When budgeting, model Slack plan seats and agent subscriptions as distinct cost lines.

Is Claude Code's 50% weekly usage-limit boost permanent?

No — and the change is now dated. Anthropic announced August 29, 2026 that the temporary 50% weekly-limit boost (active since May 13, 2026) ends on September 14, 2026, when a permanent 25% increase over the pre-promotion baseline takes effect for Pro, Max, Team, and seat-based Enterprise plans. Because 125% replaces 150%, paid users see a net 17% reduction in weekly Claude Code capacity versus today — Anthropic confirmed the number itself. Five-hour session limits, plan pricing, and API billing are unchanged. Free and consumption-based Enterprise seats are excluded. See our full breakdown of the September 14 change.

Does Cursor still use OpenAI?

Yes — through November 12, 2026. OpenAI notified SpaceX on August 28, 2026 that it intends to wind down its model-supply contract with Cursor, giving the maximum notice its contract allows. After November 12, Cursor will no longer receive OpenAI models, including the upcoming Astra model. Cursor says OpenAI models currently serve about 5% of its user traffic.

Which AI coding agent is best in 2026?

The August 2026 agent harness rankings put Claude Code first and Codex CLI second in both Machine Brief's top five and explainx.ai's top ten, with Cursor third. On the model side, Anthropic's Claude Fable 5.1 (Sept 1, 2026) is the strongest argument for the "best AI model for coding agents 2026" label: Terminal-Bench 4.0 55.8% vs Opus 5's 52.3% and GPT-5.6 Sol's 37.3%, CursorBench 3.2.0 73.4%, at $10/$50 per 1M tokens with $0.25 cache reads. The day after Fable 5.1 shipped, Google answered with Gemini 3.8 Flash (Sept 2, 2026) — $0.75/$3.75 per 1M through Dec 31 (introductory), Google claiming the top of the DeepSWE v1.1 leaderboard at a fraction of the cost of larger frontier models — so the "best" verdict now depends on whether you optimize for benchmark ceiling (Fable 5.1) or cost-per-capability (Gemini 3.8 Flash / Grok 4.6). "Best" also depends on workflow: Claude Code is the default choice for long autonomous terminal sessions, Codex CLI leads on cloud, pull-request-shaped autonomy, and Cursor holds in-editor workflows. The rankings themselves note they score hooks, subagents, and workflow depth — not onboarding friction or cost per session, which is what your budget actually feels. One 2026 factor the rankings predate: Cursor loses direct access to OpenAI models on November 12, 2026 (no Astra access either), while Anthropic is expanding Claude compute in Cursor — so re-test in-editor workflows on Claude/Gemini/Grok before standardizing on Cursor.

Claude Code vs Codex CLI: which should you use?

Both run similar frontier models and rank #1 and #2 in the August 2026 harness rankings — the difference is session management. Choose Claude Code for long autonomous sessions, terminal-native work, and the deepest hooks, subagents, and dynamic workflows, with weekly usage limits as the cost constraint. Choose Codex CLI for cloud-based, pull-request-shaped autonomy and Work mode deliverables (documents, spreadsheets, reports), with usage limits shared across Chat, Work, and Codex modes.

Which agent harness should agencies standardize on?

For the least setup friction, Claude Code or Cursor — both have the most polished onboarding and sensible permission defaults out of the box. For enterprises standardizing across many teams, Factory's Droid model or GitHub Copilot's SDK provide a platform for consistent, scoped agent configurations. For open-source, self-host, or data-residency requirements, OpenCode for provider breadth or Pi to modify the loop itself, with Aider for git-native reviewable commits. Decide on compliance and data residency, team size, model flexibility, customization depth, and budget shape.

Sources

Note: the source guide's governance section promotes TrueFoundry's own AI Gateway product; its latency/RPS claims are vendor marketing. The billing models, cost variables, and budgeting guidance cited here are independent of that.