06-reference

alphasignal gpt61 sol cost cut decisions api

2026-09-30·reference·source: AlphaSignal·by Lior Alexander
openaigpt-6.1-soldecisions-apimodel-pricinginference-costmodel-routingqwen3

Why this is in the vault

OpenAI's new Decisions API — a hosted, 150ms if-else-replacement classifier that routes app logic to an action — is a productized version of the exact hand-rolled routing pattern RDCO already runs (auto-mode classifier hard-gate, model/effort delegation pairing), and the same-day GPT-6.1 Sol cost cut extends the dated "frontier commoditizing in real time" marker started in [[2026-09-23-alphasignal-opus55-gpt6-sol-luna-pricing]].

Mapping against Ray Data Co

Concrete connection: the Decisions API is exactly the shape of two things RDCO already does by hand rather than by API call. feedback_automode_classifier_hard_gate gates deploy/production-write actions through a classifier before allowing them through; feedback_delegation_model_effort_pairing decides, per task, which model + effort tier a delegation gets. OpenAI is now selling that pattern as a primitive: hand it "a question and a list of possible answers," it picks one, on a tuned model (GPT-6 Luna), in 150ms versus 1.6s for a regular call, at API-customer pricing not yet public. The concrete question this raises for RDCO — not urgent, but worth a note for the next harness-thesis review — is whether outsourcing the classify-then-route step to a managed decision endpoint is cheaper and more maintainable than the current logic embedded in settings.json gates and delegation heuristics, or whether keeping it in-house is the right call precisely because those gates are safety-critical (deploy/production-write denial) and shouldn't depend on a third party's uptime or classification drift. Second, weaker but reinforcing: GPT-6.1 Sol's headline number — cached input down to $0.10/M tokens, a 95% discount off standard input pricing, while matching flagship Astra on coding at one-fifth the cost — is a second same-week data point (after Opus 5.5's 40% cut covered in the 09-23 note) that cost is decreasingly the binding constraint on scaling agent usage. That's the same direction as [[2026-09-29-mostlymetrics-token-cost-budgeting-frontier-convergence]]'s CFO-budgeting argument: if frontier and near-frontier pricing keeps converging downward, the planning problem shifts from "can we afford this model" to "which task actually needs the expensive one," which is the same question the delegation-pairing memory already answers manually.

Curation section

Zero deep-fetches this issue — the newsletter's own body already carries the specific benchmark/pricing numbers for both Top News items, and the Signals items don't clear the "specific hook + plausible RDCO relevance" bar beyond what's summarized above.

⚠️ Sponsorship

Three distinct paid placements this issue, none overlapping the same-issue editorial content:

The masthead "In Partnership with" slot is present but unresolved in this issue — plaintext extraction shows the label immediately followed by "Today's Author" with no legible name or link in between, consistent with the majority-unresolved pattern (clean resolutions only on 09-09/QA.tech and 09-25/Voices).

Related