Back in August I wrote about phase-scheduled activation and context pruning at 80-95% usage — mostly reverse-engineered from watching my own multi-agent pipeline blow through token budgets and guessing at thresholds that felt right. Reasonable guesses, it turns out. On September 21st AWS open-sourced Strands Harness, and buried in the release notes are the exact numbers I was guessing at: truncate tool results over 1,500 tokens, trigger summarization once context usage crosses 85%. Someone ran the benchmarks. I just ran the vibes.

That’s the real story here, not the “AWS ships another agent framework” headline. Strands Harness is a general-purpose harness on top of the existing Strands SDK, Apache 2.0, model-agnostic — Bedrock, Anthropic, OpenAI, Google, Ollama, LiteLLM, pick your poison. Across six benchmarks (ALFWorld, GAIA, WebShop, Terminal-Bench 2.1, and a couple others), it runs 28% cheaper in tokens than comparable harnesses. On Terminal-Bench 2.1 specifically, it beat Claude Code’s own harness on score while costing 77% less. Read that twice. Not “close to Claude Code for less.” Better, for a quarter of the price.

Where the money actually goes

Every agent harness has the same leak: tool results. You call a search API, a file-read, a shell command — and the raw output goes straight into the context window, verbatim, every single turn, forever, until something evicts it. A grep -r across a medium repo can dump 8,000 tokens into context for a result the model needs maybe 200 tokens of. Multiply that by a 40-turn agent loop and you’re burning real money on text nobody reads twice.

Strands Harness’s answer is almost insultingly simple: if a tool result is over 1,500 tokens, truncate it. Not a summarization pass, not an LLM call to compress it — a hard cutoff. The assumption baked in is that if a result is that large, the useful signal is near the top (error messages, first N matches, file headers) and the tail is noise. Then separately, once the whole context window crosses 85% usage, it runs an actual summarization pass to compact history rather than truncate the newest thing.

Two different problems, two different knobs. I’d been treating “big tool output” and “context getting full” as the same problem and solving both with one summarization call, which is slower and more expensive than a truncation check. Splitting them is the kind of decision that only comes from actually running the benchmarks, not from intuition.

A minimal version you can steal today

You don’t need the whole SDK to get 80% of the benefit. This is roughly what the truncation half looks like as middleware around any tool call:

MAX_TOOL_RESULT_TOKENS = 1500
CONTEXT_SUMMARIZE_THRESHOLD = 0.85

def wrap_tool_result(result: str, token_counter) -> str:
    tokens = token_counter(result)
    if tokens <= MAX_TOOL_RESULT_TOKENS:
        return result
    # keep the head, note what was cut — don't just chop silently
    head = token_counter.truncate(result, MAX_TOOL_RESULT_TOKENS)
    return f"{head}\n\n[...truncated, {tokens - MAX_TOOL_RESULT_TOKENS} more tokens omitted]"

def maybe_summarize(context, context_window_size):
    usage = context.token_count() / context_window_size
    if usage < CONTEXT_SUMMARIZE_THRESHOLD:
        return context
    return summarize_older_turns(context, keep_last_n_turns=3)

The detail that matters: wrap_tool_result runs on every single tool call, cheap and synchronous. maybe_summarize runs rarely and costs an LLM call, so it only fires when you’re actually close to the ceiling. If you flip that — summarize aggressively, truncate rarely — you pay for compression you don’t need most of the time.

Where I’d push back

I’m not going to pretend a fixed 1,500-token cutoff is universally correct. It’s a good default for shell output, file grep, API responses — anything where relevance decays with position. It’s a bad default for something like a full SQL schema dump or a long structured JSON payload where the field you need might be node 400 out of 500. AWS’s benchmarks are aggregate numbers across six tasks; your agent’s tool mix determines whether 1,500 is generous or brutal. Worth instrumenting your own truncation-hit rate before trusting the default blind.

The bigger shift is what this release signals: context engineering stopped being a craft technique individual teams reinvent and started being a benchmarked, published discipline with actual A/B numbers behind the thresholds. Six months ago “what’s your truncation limit” was a question you answered with a shrug and a number that felt okay. Now there’s a public benchmark suite you can run against your own harness and know if 1,500 tokens is actually buying you anything, or if you’re leaving performance on the table by cutting too aggressive.

If you’re running a custom agent harness in production, the homework isn’t “switch to Strands.” It’s: pull the numbers, run your own workload against both cutoffs, and see which one your tool mix actually rewards. I did. Turns out my August guess of 80% for the summarization threshold was close enough not to matter, but I was truncating tool results too conservatively — 3,000 tokens instead of 1,500. Cut it down, ran the same pipeline again, same output quality, noticeably fewer tokens per run. Small fix, free money.

Export for reading

Comments