claude-sonnet-5-5
Claude Sonnet 5.5 is Anthropic's model for well-scoped everyday work — building features, fixing bugs, and the tasks that make up most of a working day. It carries the same one-million-token window and 128,000-token output ceiling as the tier above it, reads images as well as text, and runs at Anthropic's Fast latency rating. Its adaptive thinking defaults to high effort, which is a level above what the larger model defaults to. And it drops a feature its predecessor had: it no longer receives the injected tags that let earlier Sonnet models track their own remaining context.

Claude Sonnet 5.5
Well-scoped everyday work — at a default effort higher than the tier above it.
The Default Effort Is Higher Than Opus
The comparison worth making first, because it runs against expectation.
| Claude Sonnet 5.5 | Claude Opus 5.5 | |
|---|---|---|
| Context window | 1M | 1M |
| Max output | 128K | 128K |
| Max output (Batch, beta) | 300K | 300K |
| Thinking | Adaptive | Adaptive |
| Default effort | high | medium |
| Latency | Fast | Moderate |
| Knowledge cutoff | Jun 2026 | Jun 2026 |
| Input | Text, images | Text, images |
The smaller model deliberates more by default.
Which makes sense once you invert it. The larger model's default dropped to medium because it
reaches its conclusions in fewer tokens. This one defaults to high because it needs a longer path to
arrive somewhere comparable.
So the two defaults describe the same destination reached differently. One thinks less and gets there; one thinks more and gets there.
And the latency rating still favours this model. Fast against Moderate, at a higher default
effort — which tells you the per-token speed difference is large enough to absorb the extra
deliberation.
Read the pair as a genuine choice rather than a ladder. More thinking at higher speed, or less thinking at greater depth per token. Which wins on your work is measurable and not obvious.
Well-Scoped Everyday Work
Anthropic's own positioning, and the phrasing is precise in both halves.
Well-scoped — a task with clear boundaries. A bug with a reproduction. A feature with a specification. A document with a defined output.
Everyday — the work that makes up most of a working day, at the volume that implies.
And the named strengths follow: building features and fixing bugs.
What that framing excludes is as informative as what it includes. Open-ended architectural judgment, multi-day investigations, and problems where the scope is itself the question are where the tier above is positioned.
Which is a useful division because most work is not that. A model tuned for the well-scoped majority, running fast, is the one you call constantly — and the one whose cost and latency decide whether an assistant is in the workflow or beside it.
Context Awareness Was Removed
A documented behaviour change between this model and the one it replaces, and it is easy to miss.
Earlier Sonnet models have context awareness — Anthropic inject tags that let the model track its remaining context window throughout a conversation, so it knows how much budget it has left.
Claude Sonnet 5.5 does not receive those tags.
What that changes. The previous model could notice it was running out of room and adjust — summarising earlier, wrapping up, or flagging that compaction was due. This one cannot, because it does not know.
Which moves the responsibility to you. Tracking context consumption across a long session, and deciding when to compact, is now entirely your code's job.
WINDOW = 1_048_576
COMPACT_AT = 0.80
def check_budget(prompt_tokens: int) -> None:
"""Warn before the window fills — the model will not warn you."""
used = prompt_tokens / WINDOW
if used > COMPACT_AT:
raise RuntimeError(
f"context {used:.0%} full ({prompt_tokens:,} of {WINDOW:,}) — compact before continuing"
)Check usage.prompt_tokens on every turn of a long session. It is the only signal you get.
And note that the larger model in this generation also lacks context awareness — so this is a family-level direction rather than a Sonnet-specific loss.
⚠️ Sampling Parameters May Be Rejected
Flagged rather than asserted, because the documentation covers the predecessor.
Claude Sonnet 5 introduced two hard rejections, documented as behaviour changes:
Manual extended thinking returns a 400 error. Thinking is adaptive; requesting it explicitly is rejected rather than honoured.
Setting temperature, top_p, or top_k to non-default values returns a 400 error.
That second one is the consequential one for an integration. A request setting temperature: 0.7
— entirely normal on almost every other model — fails.
This model is described as a direct upgrade from that one, which makes the behaviour likely to carry. It is not separately documented for this version, so verify it rather than assuming either way.
# One test request settles it.
try:
client.chat.completions.create(
model="anthropic/claude-sonnet-5-5",
messages=[{"role": "user", "content": "hi"}],
max_tokens=16,
temperature=0.7,
)
print("temperature accepted")
except Exception as exc:
print(f"temperature rejected: {exc}")If it is rejected, strip sampling parameters at your integration boundary rather than letting every caller discover it. A 400 on a parameter that works everywhere else is a confusing error to receive.
Thinking Blocks Are Kept and Billed
Documented behaviour across this generation.
The API keeps previous thinking blocks by default, and they count toward the context window like any other input tokens.
Two consequences on a multi-turn session.
Consistency improves. The model sees how it reasoned earlier rather than re-deriving it, which is the point of retaining them.
And input grows faster than the visible conversation. Each turn's reasoning joins the history and is charged as input on every subsequent turn.
Which compounds with the high default effort. A model defaulting to high produces more reasoning
per turn, and all of it is retained.
On a long session, watch input tokens climb rather than counting turns. The gap between the conversation you can see and the tokens you are paying for is the retained reasoning.
A Million In, 128,000 Out
| Context window | 1M tokens |
| Max output | 128K tokens |
| Max output (Batch API, beta) | 300K tokens |
The batch figure is more than double the synchronous ceiling — three hundred thousand tokens of output in one response, which makes a complete long document or a full refactor a single generation rather than an assembly job.
And the ceilings match the tier above exactly. Same window, same output limits, same batch allowance — the difference between the two models is not capacity.
⚠️ Token Counts Changed
A practical note documented for the previous generation and worth carrying forward.
Token counts for the same text run higher than on earlier Claude models, because the tokenizer changed. Anthropic's own guidance was direct: do not reuse counts measured against earlier models; recount.
And the second-order effect is the one people miss. The window is a million tokens, but each token covers less text on average — so the same window holds less text than the same number would have held on an older model.
Which means a document that fit before may not fit now, even though the advertised window did not shrink.
Recount against this model rather than trusting an estimate carried over from another.
Specifications
| Model ID | anthropic/claude-sonnet-5-5 |
| Context window | 1M tokens |
| Max output | 128K tokens |
| Max output (Batch API, beta) | 300K tokens |
| Thinking | Adaptive |
| Default effort | high |
| Comparative latency | Fast |
| Input → output | Text and images → text |
| Reliable knowledge cutoff | June 2026 |
| Training data cutoff | June 2026 |
| Context awareness | Not present |
| Status | Latest |
| Weights | Closed |
| Developer | Anthropic |
Available across Anthropic's own API and the major cloud platforms — Amazon Bedrock, Google Cloud, Microsoft Foundry, and Claude Platform on AWS.
Capabilities
| Capability | Value |
|---|---|
input_types | text, image |
output_types | text |
context_window | 1048576 |
max_output_tokens | 131072 |
reasoning | Adaptive |
default_effort | high |
streaming | Supported |
tool_calling | Supported |
structured_output | Supported |
prompt_caching | Supported |
context_awareness | Not present |
requires_prompt | Yes — text prompt required, image optional |
Using Claude Sonnet 5.5 on DEVUP AI
Base URL: https://api.devupai.com/v1 · Model ID: anthropic/claude-sonnet-5-5
Python
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["DEVUP_API_KEY"],
base_url="https://api.devupai.com/v1",
)
response = client.chat.completions.create(
model="anthropic/claude-sonnet-5-5",
messages=[
{"role": "user", "content": "Hello world!"}
],
max_tokens=1024,
)
print(response.choices[0].message.content)Node.js
import DevupAI from "devupai";
const client = new DevupAI({
apiKey: process.env.DEVUP_API_KEY,
});
async function main() {
const response = await client.chat.completions.create({
model: "anthropic/claude-sonnet-5-5",
messages: [{ role: "user", content: "Hello world!" }],
max_tokens: 1024,
});
console.log(response.choices[0].message.content);
}
main();cURL
curl -X POST "https://api.devupai.com/v1/chat/completions" \
-H "Authorization: Bearer $DEVUP_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "anthropic/claude-sonnet-5-5",
"messages": [
{ "role": "user", "content": "Hello world!" }
],
"max_tokens": 1024
}'⚠️ Send no sampling parameters until you have tested whether they are accepted. The predecessor rejected them with a 400, and a request that works on every other model in this catalogue may not work here.
Building a Feature
The named strength, and the prompt shape that suits a well-scoped model.
SYSTEM = """You are implementing a feature in an existing codebase.
Before writing code: state what you understood the requirement to be, and name anything the
specification leaves ambiguous. If something is ambiguous, ask rather than choosing.
When you write code: match the surrounding style, handle the error cases, and add tests for the
behaviour you added.
When you finish: list what you changed and what you did not, and say what you would want reviewed."""
response = client.chat.completions.create(
model="anthropic/claude-sonnet-5-5",
messages=[
{"role": "system", "content": SYSTEM},
{"role": "user", "content": f"{codebase_context}\n\nRequirement:\n{specification}"},
],
max_tokens=32768,
)"Ask rather than choosing" is what makes a well-scoped model behave well on an under-scoped task. This model is tuned for work with clear boundaries; the useful behaviour when the boundaries are unclear is to say so rather than to invent them.
And "say what you would want reviewed" turns a completed task into a reviewable one. A model that flags its own uncertain decisions saves the reviewer finding them.
Fixing a Bug
SYSTEM = """You are diagnosing a defect.
Find the root cause before proposing a fix. Quote the code that causes it and explain the condition
that triggers it.
Propose the smallest change that fixes the cause rather than the symptom. If the smallest correct fix
is larger than it looks, say why.
If the information given is not enough to identify the cause, say what you would need — a log, a
stack trace, a specific file — rather than guessing."""
response = client.chat.completions.create(
model="anthropic/claude-sonnet-5-5",
messages=[
{"role": "system", "content": SYSTEM},
{"role": "user", "content": f"{code}\n\nSymptom: {description}"},
],
max_tokens=16384,
)"The cause rather than the symptom" is the instruction that separates a fix from a patch. A condition that returns the wrong value can be worked around at the call site or corrected at the source, and only one of those stops it recurring.
And the escape hatch matters more on a bug than on a feature. A model asked to diagnose with insufficient information will produce a plausible diagnosis, and a plausible wrong diagnosis costs more than an honest request for a log.
An Agent Loop, With Your Own Budget Tracking
The pattern the missing context awareness requires.
import json
import time
WINDOW = 1_048_576
WARN_AT = 0.70
STOP_AT = 0.85
TOOLS = [
{
"type": "function",
"function": {
"name": "read_file",
"description": "Read a file relative to the repository root.",
"parameters": {
"type": "object",
"properties": {"path": {"type": "string"}},
"required": ["path"],
},
},
},
{
"type": "function",
"function": {
"name": "run_tests",
"description": "Run the test suite and return pass/fail counts with failure output.",
"parameters": {"type": "object", "properties": {}},
},
},
]
def read_file(path: str) -> dict:
"""Replace with your real, sandboxed file access."""
raise NotImplementedError
def run_tests() -> dict:
"""Replace with your real, sandboxed test runner."""
raise NotImplementedError
HANDLERS = {"read_file": read_file, "run_tests": run_tests}
session = [
{
"role": "system",
"content": (
"You are working inside a git repository. Find the root cause before changing anything, "
"and run the tests after every edit."
),
},
{"role": "user", "content": "The DZD invoice test fails on totals ending in .005. Find the cause."},
]
CEILING = 80
start = time.monotonic()
for step in range(CEILING):
response = client.chat.completions.create(
model="anthropic/claude-sonnet-5-5",
messages=session,
tools=TOOLS,
max_tokens=32768,
)
used = response.usage.prompt_tokens / WINDOW
if used > STOP_AT:
print(f"stopping at {used:.0%} of the window — compact and resume")
break
if used > WARN_AT:
print(f"warning: {used:.0%} of the window used at step {step}")
message = response.choices[0].message
session.append(message) # thinking blocks travel with it and are billed next turn
if not message.tool_calls:
print(message.content)
break
for call in message.tool_calls:
handler = HANDLERS.get(call.function.name)
if handler is None:
outcome = {"error": "unknown tool", "name": call.function.name}
else:
try:
outcome = handler(**json.loads(call.function.arguments or "{}"))
except Exception as exc:
outcome = {"error": type(exc).__name__, "detail": str(exc)}
session.append({"role": "tool", "tool_call_id": call.id, "content": json.dumps(outcome)})
else:
print(f"Reached the {CEILING}-step ceiling.")
print(f"{step + 1} steps, {(time.monotonic() - start) / 60:.1f} min")The budget check is the part that earlier Sonnet models did for themselves. This one does not track its remaining window, so the loop has to.
Two thresholds rather than one. A warning gives you a signal in the logs; a stop prevents the run from failing at the ceiling with nothing recoverable.
And appending the message object as returned keeps the thinking blocks with it — which is what preserves consistency across the session, and what makes the input column climb.
Images
Text and images in, text out — the same input modalities as the tier above.
import base64
from pathlib import Path
encoded = base64.b64encode(Path("screenshot.png").read_bytes()).decode("utf-8")
response = client.chat.completions.create(
model="anthropic/claude-sonnet-5-5",
messages=[
{
"role": "user",
"content": [
{"type": "image_url", "image_url": {"url": f"data:image/png;base64,{encoded}"}},
{
"type": "text",
"text": (
"Describe exactly what is rendered incorrectly, referring to elements by "
"their visible labels. Do not speculate about the cause in code."
),
},
],
}
],
max_tokens=16384,
)Separating observation from diagnosis keeps the output verifiable. What is on screen can be checked against the image; why is a hypothesis about code the model has not read.
Send images at full resolution. Downscaling first discards detail the model would use.
Long-Form Generation
128,000 tokens synchronously, and up to 300,000 through the batch interface.
response = client.chat.completions.create(
model="anthropic/claude-sonnet-5-5",
messages=[
{
"role": "system",
"content": (
"Produce the complete document requested. Use headings. Do not stop early, do not "
"summarise sections you have not written, and do not add meta-commentary."
),
},
{"role": "user", "content": f"{source_material}\n\nWrite the full specification."},
],
max_tokens=100_000,
)
choice = response.choices[0]
if choice.finish_reason == "length":
raise ValueError(f"reached {response.usage.completion_tokens:,} tokens without finishing")The budget covers reasoning and document together. Thinking is adaptive and defaults to high, so
a hundred thousand tokens is not a hundred thousand tokens of prose.
Lower the effort on a long generation if your path supports setting it. Writing a document from
supplied material is not the kind of work high effort was for, and the tokens it spends are tokens
the document does not get.
Comparing Against the Tier Above
Both models share the window, the output ceilings, the input modalities, and the knowledge cutoff. What differs is speed, default effort, and what each does with a token.
import time
MODELS = ["anthropic/claude-sonnet-5-5", "anthropic/claude-opus-5-5"]
def compare(system: str, prompt: str) -> None:
"""Run the same task on both and report tokens and wall-clock time."""
for model in MODELS:
start = time.monotonic()
response = client.chat.completions.create(
model=model,
messages=[
{"role": "system", "content": system},
{"role": "user", "content": prompt},
],
max_tokens=32768,
)
elapsed = time.monotonic() - start
print(f"{model:<32} {response.usage.completion_tokens:>7,} tokens {elapsed:>6.1f}s")
print(response.choices[0].message.content[:400])
print("---")Compare on well-scoped tasks first — that is where this model is positioned, and where the answer is least obvious.
Watch wall-clock time, not only tokens. A Fast model at high effort and a Moderate model at
medium can produce similar token counts at very different latencies, and on an interactive path
latency is what the user experiences.
Twenty of your real tasks settles it better than any positioning statement, including Anthropic's own.
Where It Fits
Feature work and bug fixing, named directly as its strengths.
Well-scoped everyday tasks at the volume that description implies — the model you call constantly rather than reserve.
Interactive paths, where the Fast latency rating is the product.
Long-context analysis at a million tokens, with the same window as the tier above.
Long-form generation at 128,000 tokens, or 300,000 through batch.
Document and screenshot reading.
Agent loops, with your own context tracking in place of the awareness earlier Sonnet models had.
Less suited to open-ended problems where the scope is itself the question, and to work where the tier above's depth per token is what you are paying for.
Not for audio or video. Text and images in, text out.
Practical Notes
Test whether sampling parameters are accepted before letting callers set them.
Track context consumption yourself — this model does not track its own.
Watch input tokens climb across a session; retained thinking is charged as input every turn.
Recount tokens against this model rather than reusing counts from an older one.
Size max_tokens for reasoning plus answer; the default effort is high.
Lower the effort on long-form generation if your path allows it.
Check finish_reason on every request.
Instruct the model to ask rather than choose when a requirement is ambiguous.
Compare against the tier above on twenty well-scoped tasks, watching latency as well as tokens.
Limitations
No context awareness. Earlier Sonnet models tracked their remaining window; this one does not, and that tracking is now your code's job.
Sampling parameters may be rejected. Documented as a 400 on the predecessor and not separately documented here — test before relying on either behaviour.
Manual extended thinking was rejected on the predecessor. Thinking is adaptive; requesting it explicitly is not how it is controlled.
Retained thinking grows input tokens on every turn, and the high default effort produces more of it per turn.
Token counts run higher than on older Claude models, so the same window holds less text than the figure suggests.
Positioned for well-scoped work. Anthropic's own framing, and the tier above exists for what falls outside it.
Knowledge ends June 2026. Recent by current standards, and a fixed point — ground anything time-sensitive.
Text and images in, text out. No audio, no video, no image generation.
Closed weights. API access only, with no architecture published and no self-hosted option.