A task budget is a property of the whole run, not a line of instruction copied into each prompt. Copying $5 remaining to ten independent branches creates ten claims on the same money. This is a distributed resource allocation problem. A graph recursion limit can cap one kind of loop, and frameworks such as LangGraph expose graph step limits, but a step count is not a cost or token budget. Model prices, tool fees, context size and output length differ. The runtime must account for those costs across the parent and all descendants.

Before spawning, I would reserve from one authoritative run ledger. If the parent has $5, it can allocate, say, five children $0.70 each, keep $1.50 for synthesis and retries, and decline further branches. The numbers are a policy example, not a fixed formula. A child receives a scoped budget ID and cannot create more capacity by spawning its own children. At admission, reserve an upper bound or a conservative allowance for an in-flight call. On completion, reconcile actual metered cost and release unused reservation. If the provider's exact bill arrives later, keep an auditable estimate and a final settlement path. Do not let dozens of concurrent calls all pass a check against the same unreserved balance.

There are two limits to distinguish. A soft cost target can stop new work once observed spending reaches it, but cannot guarantee a hard cap on already admitted calls. A hard cap needs admission limits on maximum input and output, upper-bound pricing for the chosen model and tools, and control over retries. Some external tools have unpredictable downstream charges, so the product may need a separate quota or conservative ceiling. Cancellation after the budget is exhausted should release reservations only after in-flight work has finished or been confirmed canceled. A retry of the same operation should not double-count the reservation, nor should it evade the cap with a new branch ID.

I would test ten subagents all requesting the last dollar at once, a child that spawns another child, provider timeouts with uncertain usage, and a paused run that resumes after pricing or quota policy changes. Show per-run spend, committed and reserved amounts, and why a branch was refused. The multi-agent demo is better, but production pays for it asks whether multi-agent quality is worth production cost. This question asks how the agreed limit remains true when the execution graph fans out.