5  Chapter 4: Dynamic Allocation — Tool Calling

5.1 Tool Calls as Memory Allocations

If the messages array is memory, then every tool call is a malloc(). Each tool interaction adds at minimum three entries to the array:

  1. The assistant’s message containing the tool call (the function name + arguments)
  2. The tool result (returned by your application)
  3. The assistant’s response interpreting the result

The token cost of a single tool call cycle breaks down like this:

Assistant decides to call a tool:
  "I'll read the file src/auth.ts"
  → ~50-100 tokens (reasoning + tool_use block)

Tool call payload:
  {"name": "read_file", "input": {"path": "src/auth.ts"}}
  → ~30-50 tokens

Tool result (the file contents):
  → 200-5,000+ tokens (depends entirely on what comes back)

Assistant processes the result:
  "The auth module exports three functions..."
  → 100-500 tokens

Total for one tool call: 400-6,000+ tokens. A single file read can consume as many tokens as the entire system prompt.

To see what this looks like in the actual API payload, here’s the messages array after one tool call cycle:

[
  {"role": "system", "content": "You are a coding assistant..."},
  {"role": "user", "content": "What's in config.json?"},
  {"role": "assistant", "content": [
    {"type": "tool_use", "id": "call_1", "name": "read_file",
     "input": {"path": "config.json"}}
  ]},
  {"role": "user", "content": [
    {"type": "tool_result", "tool_use_id": "call_1",
     "content": "{\n  \"db_host\": \"localhost\",\n  \"db_port\": 5432,\n  \"auth\": {\"provider\": \"oauth2\", \"timeout\": 3600}\n}"}
  ]},
  {"role": "assistant", "content": "The config has database settings (localhost:5432) and OAuth2 auth with a 3600s timeout."}
]

Five messages from one question. And every one of them stays in the array for every subsequent API call. The JSON config content — which the model already interpreted — is re-sent and re-processed on every future turn.

This creates an accumulation problem: tool results are append-only. Once a 3,000-token file is read into context, it stays there for the rest of the session. Read 10 files? That’s potentially 30,000 tokens of file content that persists in your context forever (within that session).

To visualize how the array grows during a typical interaction:

[system]                            → 1,500 tokens  (cumulative: 1,500)
[harness]                           → 3,000 tokens  (cumulative: 4,500)
[user: "Fix the login bug"]         → 20 tokens     (cumulative: 4,520)
[assistant: "Let me look..."]       → 100 tokens    (cumulative: 4,620)
[tool_call: search_code]            → 80 tokens     (cumulative: 4,700)
[tool_result: 15 matches]           → 2,000 tokens  (cumulative: 6,700)
[assistant: "Found it, reading..."] → 150 tokens    (cumulative: 6,850)
[tool_call: read_file]              → 50 tokens     (cumulative: 6,900)
[tool_result: file contents]        → 4,000 tokens  (cumulative: 10,900)
[assistant: "I see the bug..."]     → 300 tokens    (cumulative: 11,200)
[tool_call: edit_file]              → 200 tokens    (cumulative: 11,400)
[tool_result: "success"]            → 30 tokens     (cumulative: 11,430)
[assistant: "Fixed. Let me test"]   → 100 tokens    (cumulative: 11,530)
[tool_call: run_tests]              → 50 tokens     (cumulative: 11,580)
[tool_result: test output]          → 8,000 tokens  (cumulative: 19,580)
[assistant: "All tests pass."]      → 80 tokens     (cumulative: 19,660)

In just one bug fix cycle: ~15,000 tokens of dynamic context consumed. That’s ~22% of a 68K effective budget (from Chapter 2’s calculation) — gone in one task.

5.2 A Real Agent Session

Consider what happens when an agent handles a more complex request — say, “Add user authentication to this Express app.” A realistic multi-task session unfolds in four phases.

Phase 1: Understanding covers exploring the codebase. The agent searches for existing auth patterns (~2,000 tokens), reads package.json (~800 tokens), reads app.ts (~3,000 tokens), and reads existing middleware (~2,500 tokens). Running total: ~8,300 new tokens.

Phase 2: Implementation is where code gets written. Creating auth middleware costs ~500 tokens (tool call + result), creating the user model ~400 tokens, updating routes ~600 tokens, and creating the login endpoint ~500 tokens. Running total: ~10,300 new tokens.

Phase 3: Testing is where context pressure really builds. Running tests produces ~5,000 tokens of output. Reading error details adds ~2,000 tokens. A fix costs ~300 tokens, but then re-running tests dumps another ~5,000 tokens. Another fix at ~300 tokens, another test run at ~5,000 tokens. Running total: ~28,200 new tokens.

Phase 4: Cleanup rounds out the session with the model now deep in context. Updating the README adds ~500 tokens, and a final test run produces another ~5,000 tokens. Running total: ~33,700 new tokens.

The total context at end is ~11,500 (fixed allocations from Chapter 2) + 33,700 (dynamic) = ~45,200 tokens. Still in the smart zone for a 200K model, but only for a single feature addition. Two or three more tasks in the same session, and you’re in the dumb zone.

Notice where context pressure hits hardest: test output. Testing alone consumed ~20,000 tokens — more than half the session’s dynamic context. Each test run dumps its full output into the array. This is the #1 context pressure point in agent sessions. (Chapter 5 introduces the solution: sub-agents for test running.)

5.2.1 Signs of session degradation

When a session starts running long, you’ll see predictable symptoms:

  1. The agent re-reads files it already read
  2. It forgets constraints you stated earlier
  3. Tool calls become less precise (searching broadly instead of targeted reads)
  4. It starts apologizing and “trying again” without changing approach
  5. Code quality drops — more bugs, less coherent architecture

When you notice these signs, the most productive action is often to start a new session, not to “remind” the agent. Reminding adds more tokens to an already-stressed context. Starting fresh (with a well-designed spec) gives the model a clean smart zone to work in. This is the bridge to Chapter 7.

5.3 Key Takeaways

  • Every tool call is a memory allocation — it grows the context permanently within a session.
  • A single file read can cost 3,000-5,000 tokens. Test output is the #1 context hog.
  • A typical agent session for one feature can consume 30,000-50,000 tokens of dynamic context.
  • Context pressure shows up as degraded behavior: repetition, forgotten instructions, imprecise tool use.
  • The solution isn’t “more context” — it’s fresh context with the right state carried forward.

5.4 References

  • Anthropic. “Tool Use (Function Calling) Documentation.” https://docs.anthropic.com/en/docs/build-with-claude/tool-use
  • OpenAI. “Function Calling Guide.” https://platform.openai.com/docs/guides/function-calling
  • Schick, T., et al. (2023). “Toolformer: Language Models Can Teach Themselves to Use Tools.” arXiv:2302.04761.
  • Patil, S., et al. (2023). “Gorilla: Large Language Model Connected with Massive APIs.” arXiv:2305.15334.