Skip to Content
DeerFlow

Limits, Budgets, and Capacity

A subagent runs its own agent loop, so it needs the same backstops the Lead Agent has. This chapter puts every limit in one table and then explains what each one looks like when it fires.

Overview

LimitConfig key / context keyDefaultRangeWhen it fires
Per-response concurrencyrequest context max_concurrent_subagentssubagent_runtime.max_running (3)1 to 64, capped by subagent_runtime.max_runningExcess task calls are dropped and only logged
Per-run totalsubagents.max_total_per_run; request context max_total_subagents overrides it61 to 50Excess calls are dropped. If the total was already used up before the response, [SUBAGENT LIMIT REACHED] ... is appended to the assistant message and the run’s stop_reason becomes subagent_limit_capped; a response that only crosses the cap is trimmed with a log line
Turnssubagents.max_turns (built-ins only), subagents.agents.<name>.max_turnsgeneral-purpose 150, bash 60, custom 50at least 1Partial result kept and marked turn_capped; failed when no usable text exists
Timeoutsubagents.timeout_seconds (built-ins only), subagents.agents.<name>.timeout_secondsbuilt-ins 1800 s, custom and managed 900 sat least 1Status timed_out
Polling attemptsDerived from the timeout: (timeout_seconds + 60) / 5 polls, every 5 secondsfollows the timeoutStatus polling_timed_out; the runtime requests cancellation and schedules deferred cleanup
Token budgetsubagents.token_budget, subagents.agents.<name>.token_budgetenabled; max_tokens 2,000,000, or 1,000,000 when summarization is on; warn_threshold 0.7max_tokens at least 1000The in-flight turn is capped and forced to finish: completed + token_capped
Loop detectionloop_detection (shared with the Lead Agent)enabled; warn_threshold 3, hard_limit 5, window_size 20, tool_freq_warn 30, tool_freq_hard_limit 50completed + loop_capped
Summarizationsummarization (shared with the Lead Agent)example config: enabled, trigger at 32,000 tokens, keep the last 10 messagesSystem prompt and latest user message survive; the summary goes into summary_text and is re-injected every call
Process-wide capacitysubagent_runtime.max_running31 to 64When full, queue or reject per admission_policy
Queue boundsubagent_runtime.max_queued640 to 10,000Rejected when the queue is full: Subagent execution capacity is full (3 running, 64 queued)
Admission policysubagent_runtime.admission_policyqueuequeue / rejectWith reject, a full slot set fails immediately
Queue timeoutsubagent_runtime.queue_timeout_seconds3001 to 86,400Timed out after 300s waiting for a subagent execution slot, status failed
AIO shell sessionssandbox.environment.MAX_SHELL_SESSIONSimage default 10; auto-set to max_running + 1 when needednot below max_running + 1Startup error when too low; see Sandbox and Isolation

The four subagent_runtime fields are frozen at Gateway startup and need a restart to change. They bound ordinary task calls and durable batches alike. Queued delegations wait asynchronously without holding an execution thread. A capacity rejection or queue timeout becomes a failed result for an ordinary task; durable batch items are requeued instead.

Concurrency and total

SubagentLimitMiddleware is installed on the Lead Agent chain only and enforces two gates:

  • Per-response concurrency: how many task calls one model response may contain. It defaults to the process capacity max_running (3 by default) and is capped by it. Excess calls are dropped with no text appended, only a log line. The HARD LIMITS line in the prompt tells the model the number.
  • Per-run total: the cumulative number of delegations in one run, default 6, which is two full batches at the default concurrency. Only ledger entries tagged with the current run_id count; without a run_id the middleware logs a warning and counts the whole thread. Calls beyond the remaining total are removed. When one response merely crosses the cap, the trim is only logged. Once the total was already exhausted before a response, all of its task calls are removed, the run’s stop_reason becomes subagent_limit_capped, and the assistant message gets this appended:
[SUBAGENT LIMIT REACHED] The subagent delegation limit for this run has been reached. Continue using the subagent results already collected, execute remaining simple work directly, or summarize the remaining work instead of launching more subagents.

Both gates depend on the ledger, which DurableContextMiddleware writes. Graphs built directly with create_deerflow_agent now include it automatically; see Developers and Integration.

batch_task does not count against the per-run total; it has its own max_live_items and max_running_items.

Turns and timeouts

max_turns is the operator-facing notion of a turn: one model call plus the tools it runs. The runtime converts it into LangGraph’s super-step budget:

recursion_limit = max_turns × (nodes per turn) + (one-time nodes per invocation)

Nodes per turn is the number of before_model / after_model hooks implemented on the middleware chain plus 2 (the model node and the tools node); one-time nodes is the number of before_agent / after_agent hooks. Configuring 150 turns therefore really yields 150 turns, regardless of how many middlewares are installed.

When the turns run out the executor catches GraphRecursionError, keeps the last assistant text as a partial result, and marks it turn_capped.

The timeout is wall-clock. Once it elapses the subagent is cancelled with status timed_out. Turns and timeout are independent axes: when you raise max_turns, usually raise timeout_seconds too, or the failure merely moves from turns to timeout.

Runaway guards

The subagent middleware chain mirrors three Lead Agent guards:

  • Loop detection (LoopDetectionMiddleware): breaks a subagent that repeats the same tool call without progress. Subagents have no task, so only the tool-loop heuristic can fire. A hard stop marks the result completed + subagent_stop_reason=loop_capped. Controlled by the loop_detection config, with tool_freq_overrides for per-tool thresholds.
  • Token budget (TokenBudgetMiddleware): tracks the run’s cumulative tokens against subagents.token_budget. At the hard-stop threshold it strips the current turn’s tool calls and forces a final answer, marking the result completed + token_capped. Reaching warn_threshold injects a one-time budget warning into the subagent’s next model call and logs it at INFO level.
  • Summarization (DeerFlowSummarizationMiddleware): compacts long subagent transcripts under the same summarization.enabled switch as the Lead Agent, and by default summarizes with the subagent’s own model. Compaction preserves the subagent’s system prompt (role, report contract, acceptance note, skill index, deferred tool catalog) and the latest user message. DurableContextMiddleware sits before summarization and re-injects summary_text into later requests, so a compacted history never starts with an assistant message.

The default token ceiling is coupled to the summarization switch: 1,000,000 when compaction is on, 2,000,000 when off. An explicit subagents.token_budget.max_tokens (global or per agent) always wins, so flipping summarization never silently changes a value you pinned.

Configuration example

subagent_runtime: # restart required; shared by task and batches max_running: 3 max_queued: 64 admission_policy: queue # or reject queue_timeout_seconds: 300 subagents: timeout_seconds: 1800 # default timeout for built-in subagents # max_turns: 120 # global turn override for built-ins; unset keeps 150 / 60 max_total_per_run: 6 # delegations per run, 1 to 50 token_budget: enabled: true max_tokens: 2000000 warn_threshold: 0.7 agents: general-purpose: timeout_seconds: 2700 # 45 minutes for deep research max_turns: 250 token_budget: max_tokens: 3000000 bash: timeout_seconds: 300 max_turns: 80

Per-agent overrides beat global values. The global timeout_seconds and max_turns apply to built-in subagents only; custom and managed subagents have their own defaults, so change them through agents.<name>.

To allow strictly one subagent at a time, set the request context’s max_concurrent_subagents to 1. The floor is 1; it is not bumped to 2.