gpt-6.1-sol
GPT-6.1 Sol adds two things to the tier below OpenAI's flagship, and takes one away. It reads images, which the model it replaces did not. Its knowledge runs ten days later. And it no longer accepts the lowest reasoning settings — none and minimal are rejected, which makes deliberation mandatory and breaks any integration written against its predecessor's full range. A 1,050,000-token window with 922,000 available for input and 128,000 for output, five effort levels from low to max, and OpenAI's own documentation telling you to compare it against the flagship on your own tasks rather than take the positioning on trust.

GPT-6.1 Sol
Near-flagship capability a tier down — with one setting removed that will break your migration.
Released 29 September 2026 — model page
none and minimal Are Rejected
The first thing to check before pointing existing code at this model, stated plainly in OpenAI's own documentation:
reasoning.effortsupportslow,medium(default),high,xhigh, andmax. Thenoneandminimalreasoning efforts are not supported.
The model it replaces accepted none. This one does not.
Which means reasoning is mandatory here. There is no configuration that skips deliberation
entirely — low is the floor, and low still thinks.
Three consequences.
Existing requests will fail. Code that sets reasoning.effort: "none" — a reasonable setting for
classification, extraction, or formatting on the previous model — is rejected rather than silently
upgraded.
Your fast path needs rebuilding. Whatever you were routing to none now goes to low, or to a
different model.
And max_tokens must cover reasoning plus answer, always. On the previous model, none gave you
a budget that was entirely answer. That configuration no longer exists.
Audit your request builders before switching. This is a one-line change to find and a silent production failure if you miss it.
Vision Is New Here
Input: text and images. Output: text.
The model this replaces was text-only. Adding image input inside a point release is unusual — it is the kind of capability that normally waits for a whole-number generation.
What it enables at this tier. Document extraction, screenshot analysis, chart reading, and computer use — all without reaching for the flagship above it.
And computer use is named directly in OpenAI's positioning: complex coding, computer use, and professional work. A model driving an interface has to read the screen before it acts, which is vision and decision in one capability.
922,000 In, 128,000 Out
Three figures, and the middle one is a separate cap rather than a remainder.
| Context window | 1,050,000 tokens |
| Maximum input | 922,000 tokens |
| Maximum output | 128,000 tokens |
922,000 is a hard ceiling on input, not what remains after output. You cannot send a million tokens and ask for a short answer.
128,000 is a hard ceiling on output, not a share you enlarge by sending less.
They sum to the headline figure, which makes the total easy to satisfy and the individual caps easy to violate. A request of 950,000 input and 50,000 output comes in under 1,050,000 and still fails.
Check each separately.
OpenAI Tell You to Compare
A sentence in their own documentation worth reading twice:
Compare it with Astra on your tasks to assess the tradeoff between quality and cost.
A vendor instructing you to test rather than trust the positioning.
Which is the right advice and an unusual place to find it. "Near-flagship performance at a lower cost" is a claim about an average across tasks. Whether it holds on yours is a different question, and the only way to answer it is to run both.
The comparison is cheap. Same interface, same request shape, one changed model identifier. Twenty of your real tasks through each will tell you more than any published figure.
Specifications
| Model ID | openai/gpt-6.1-sol |
| Context window | 1,050,000 tokens |
| Maximum input | 922,000 tokens |
| Maximum output | 128,000 tokens |
| Reasoning | low, medium (default), high, xhigh, max |
| Not supported | none, minimal |
| Reasoning quality | Highest |
| Speed | Fast |
| Input | Text, image |
| Output | Text |
| Knowledge cutoff | 30 April 2026 |
| Fine-tuning | Listed as supported |
| Weights | Closed |
| Released | 29 September 2026 |
| Developer | OpenAI |
A single snapshot. gpt-6.1-sol is a dotted version rather than a dated checkpoint — there is no
gpt-6.1-sol-2026-09-29 alongside it.
A priority-tier variant exists for paths where latency outweighs everything else.
Tools through the Responses API: web search, image generation, code interpreter, and a hosted shell.
Capabilities
| Capability | Value |
|---|---|
input_types | text, image |
output_types | text |
context_window | 1048576 |
max_input_tokens | ~922,000 |
max_output_tokens | ~128,000 |
reasoning | Mandatory — low through max |
reasoning_disable | Not supported |
streaming | Supported |
tool_calling | Supported |
structured_output | Supported |
fine_tuning | Listed |
requires_prompt | Yes — text prompt required, image optional |
Ten Days of Knowledge
A small detail with a real consequence.
| Model | Knowledge cutoff |
|---|---|
| The model it replaces | 20 April 2026 |
| GPT-6.1 Sol | 30 April 2026 |
Ten days further forward.
Which sounds trivial and is not, on a fast-moving topic. A library release, a regulatory change, or a product launch inside that window is something one model knows about and the other does not.
And it is still five months back. Ground anything time-sensitive rather than relying on either.
Using GPT-6.1 Sol on DEVUP AI
Base URL: https://api.devupai.com/v1 · Model ID: openai/gpt-6.1-sol
Python
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["DEVUP_API_KEY"],
base_url="https://api.devupai.com/v1",
)
response = client.chat.completions.create(
model="openai/gpt-6.1-sol",
messages=[
{"role": "user", "content": "Hello world!"}
],
max_tokens=1024,
)
print(response.choices[0].message.content)Node.js
import DevupAI from "devupai";
const client = new DevupAI({
apiKey: process.env.DEVUP_API_KEY,
});
async function main() {
const response = await client.chat.completions.create({
model: "openai/gpt-6.1-sol",
messages: [{ role: "user", content: "Hello world!" }],
max_tokens: 1024,
});
console.log(response.choices[0].message.content);
}
main();cURL
curl -X POST "https://api.devupai.com/v1/chat/completions" \
-H "Authorization: Bearer $DEVUP_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "openai/gpt-6.1-sol",
"messages": [
{ "role": "user", "content": "Hello world!" }
],
"max_tokens": 1024
}'⚠️ How reasoning effort is set varies by platform. Confirm the parameter name and accepted values with a test request before building routing logic — and confirm what happens when an unsupported value is sent.
Guarding Against the Removed Settings
The wrapper worth writing before you migrate.
SUPPORTED_EFFORTS = {"low", "medium", "high", "xhigh", "max"}
REMOVED_EFFORTS = {"none", "minimal"}
def ask(prompt: str, *, effort: str = "medium", want: int = 8192) -> str:
"""Send a request, rejecting effort levels this model does not accept."""
if effort in REMOVED_EFFORTS:
raise ValueError(
f"'{effort}' is not supported on gpt-6.1-sol — use 'low' as the floor, "
"or route this request to a model that answers without reasoning"
)
if effort not in SUPPORTED_EFFORTS:
raise ValueError(f"unknown effort '{effort}'; expected one of {sorted(SUPPORTED_EFFORTS)}")
response = client.chat.completions.create(
model="openai/gpt-6.1-sol",
messages=[{"role": "user", "content": prompt}],
max_tokens=want,
extra_body={"reasoning_effort": effort},
)
choice = response.choices[0]
if choice.finish_reason == "length":
raise ValueError(
f"truncated at {response.usage.completion_tokens:,} tokens — "
f"raise max_tokens or lower the effort from '{effort}'"
)
return choice.message.contentFailing locally with a clear message beats failing upstream with a generic one. A migration that hits this will hit it many times, and the error naming the remedy saves the second hour.
And the error suggests routing rather than only substituting. A request that genuinely wanted no
reasoning — a classification, a format conversion — may belong on a different model entirely rather
than at low here.
What Used to Be none
The migration question this release forces, and it has two honest answers.
Route it to low here. Simplest, and it costs a reasoning pass you were not paying for. On a
high-volume classification path, that is a real change in latency and tokens.
Or route it elsewhere. Work that had no reasoning to do — extraction against a fixed schema,
routing by keyword, format conversion — was correctly assigned to none, and the lower tier in this
family is built for exactly that shape.
def route(task_kind: str) -> tuple[str, str]:
"""Return (model, effort) for a task, respecting the removed settings."""
if task_kind in ("classify", "extract", "format"):
# These wanted no reasoning. This model cannot provide that.
return ("<the lower tier in this family>", "none")
if task_kind in ("summarise", "answer"):
return ("openai/gpt-6.1-sol", "low")
return ("openai/gpt-6.1-sol", "medium")Do not route everything to low by reflex. The removed setting existed because some work does not
need thinking, and that work did not change when the model did.
Images and Computer Use
New at this tier, and named in the positioning.
import base64
from pathlib import Path
encoded = base64.b64encode(Path("screenshot.png").read_bytes()).decode("utf-8")
response = client.chat.completions.create(
model="openai/gpt-6.1-sol",
messages=[
{
"role": "user",
"content": [
{"type": "image_url", "image_url": {"url": f"data:image/png;base64,{encoded}"}},
{
"type": "text",
"text": (
"Describe exactly what is rendered incorrectly on this screen, referring to "
"elements by their visible labels. Do not speculate about the cause in code."
),
},
],
}
],
max_tokens=16384,
)Separating observation from diagnosis keeps the output checkable. What is on screen is verifiable from the image; why is a hypothesis about code the model has not read.
Send images at full resolution. Downscaling before upload discards detail the model would use.
Long-Corpus Work
922,000 tokens of input, with the output cap kept separate.
from pathlib import Path
REPO = Path("src")
sources = "\n\n".join(
f"=== {path.relative_to(REPO.parent)} ===\n{path.read_text(encoding='utf-8')}"
for path in sorted(REPO.rglob("*.py"))
)
response = client.chat.completions.create(
model="openai/gpt-6.1-sol",
messages=[
{
"role": "system",
"content": (
"You are auditing a codebase. Identify every path where a database write can occur "
"outside a transaction. Name the file, the function, and the call chain that reaches "
"it. Report nothing you cannot trace."
),
},
{"role": "user", "content": sources},
],
max_tokens=32768,
extra_body={"reasoning_effort": "high"},
)
print(f"input: {response.usage.prompt_tokens:,} of ~922,000")Asking for the call chain rather than the line is what uses a very long window rather than a search.
And the budget covers reasoning too. At high effort on a large corpus, thirty-two thousand
output tokens may be mostly trace — size accordingly and check finish_reason.
Long-Form Generation
128,000 tokens of output, shared with mandatory reasoning.
response = client.chat.completions.create(
model="openai/gpt-6.1-sol",
messages=[
{
"role": "system",
"content": (
"Produce the complete document requested. Use headings. Do not stop early, do not "
"summarise sections you have not written, and do not add meta-commentary."
),
},
{"role": "user", "content": f"{source_material}\n\nWrite the full specification."},
],
max_tokens=100_000,
extra_body={"reasoning_effort": "medium"},
)
choice = response.choices[0]
if choice.finish_reason == "length":
raise ValueError(f"reached {response.usage.completion_tokens:,} tokens without finishing")medium rather than max on a long generation. Deliberation and document share one budget, and
at the top effort levels a substantial share goes to reasoning you will not read.
"Do not stop early" earns its place. Models frequently wind down before their budget is exhausted, and an explicit instruction to complete the whole task counters it.
Validating Both Caps
MAX_INPUT = 922_000
MAX_OUTPUT = 128_000
def check_request(estimated_input: int, want_output: int) -> int:
"""Validate against both caps separately, not against their sum."""
if estimated_input > MAX_INPUT:
raise ValueError(
f"input of ~{estimated_input:,} tokens exceeds the {MAX_INPUT:,} input cap"
)
return min(want_output, MAX_OUTPUT)Two caps, checked separately. Satisfying their sum is not sufficient.
Where It Fits
Complex coding, named first in OpenAI's own positioning.
Computer use, which vision at this tier now makes possible without reaching for the flagship.
Professional and knowledge work — documents, analysis, structured output.
Long-corpus analysis at 922,000 tokens of input.
Long-form generation at 128,000 tokens of output.
Document and screenshot reading, new to this tier.
As a candidate against the flagship, which is what OpenAI's own documentation asks you to test.
Not for work that needs no reasoning at all. none and minimal are rejected; that work belongs
on the tier below.
Not for audio or video. Text and images in, text out.
Not for self-hosting. Closed weights.
Practical Notes
Audit your request builders for none and minimal before migrating. They are rejected.
Decide deliberately what used to be none — low here, or a different model.
Start at medium, the documented default.
Size max_tokens for reasoning plus answer. There is no configuration without reasoning.
Check both caps separately — 922,000 input and 128,000 output.
Check finish_reason on every request.
Send images at full resolution, and separate observation from diagnosis.
Compare against the flagship on twenty of your own tasks. OpenAI's documentation asks you to.
Ground time-sensitive work; the cutoff is April 2026.
Limitations
none and minimal are not supported. Reasoning is mandatory, and existing code using those
settings is rejected rather than adapted.
Two separate caps, not one shared budget — 922,000 input and 128,000 output, each enforced on its own.
Knowledge ends 30 April 2026 — ten days later than the model it replaces, and five months before now.
Text and images in, text out. No audio, no video, no image generation from the model itself.
Closed weights. API access only, with no architecture published and no self-hosted option.
Positioned as near-flagship, not flagship. OpenAI's own documentation frames it as a trade-off to assess rather than a replacement to assume.
Released 29 September 2026. Independent evaluation, tooling support, and catalogue integrations are all still catching up — several agent frameworks did not recognise the identifier on release day, because the dotted version does not match their existing family prefixes.
A single snapshot. No dated checkpoint to pin against, so behaviour changes arrive under the same identifier.