Why this is in the vault
OpenAI's new Decisions API — a hosted, 150ms if-else-replacement classifier that routes app logic to an action — is a productized version of the exact hand-rolled routing pattern RDCO already runs (auto-mode classifier hard-gate, model/effort delegation pairing), and the same-day GPT-6.1 Sol cost cut extends the dated "frontier commoditizing in real time" marker started in [[2026-09-23-alphasignal-opus55-gpt6-sol-luna-pricing]].
Mapping against Ray Data Co
Concrete connection: the Decisions API is exactly the shape of two things RDCO already does by hand rather than by API call. feedback_automode_classifier_hard_gate gates deploy/production-write actions through a classifier before allowing them through; feedback_delegation_model_effort_pairing decides, per task, which model + effort tier a delegation gets. OpenAI is now selling that pattern as a primitive: hand it "a question and a list of possible answers," it picks one, on a tuned model (GPT-6 Luna), in 150ms versus 1.6s for a regular call, at API-customer pricing not yet public. The concrete question this raises for RDCO — not urgent, but worth a note for the next harness-thesis review — is whether outsourcing the classify-then-route step to a managed decision endpoint is cheaper and more maintainable than the current logic embedded in settings.json gates and delegation heuristics, or whether keeping it in-house is the right call precisely because those gates are safety-critical (deploy/production-write denial) and shouldn't depend on a third party's uptime or classification drift. Second, weaker but reinforcing: GPT-6.1 Sol's headline number — cached input down to $0.10/M tokens, a 95% discount off standard input pricing, while matching flagship Astra on coding at one-fifth the cost — is a second same-week data point (after Opus 5.5's 40% cut covered in the 09-23 note) that cost is decreasingly the binding constraint on scaling agent usage. That's the same direction as [[2026-09-29-mostlymetrics-token-cost-budgeting-frontier-convergence]]'s CFO-budgeting argument: if frontier and near-frontier pricing keeps converging downward, the planning problem shifts from "can we afford this model" to "which task actually needs the expensive one," which is the same question the delegation-pairing memory already answers manually.
Curation section
- OpenAI ships GPT-6.1 Sol at 80% less cost with near-flagship performance (23,168 likes) — closes the gap to Astra on coding (DeepSWE: $0.65/task vs Astra's $3.92), computer use (OSWorld: within 2 points at one-seventh the cost), and Terminal-Bench ($5.47/task vs $23.80). Cached input now $0.10/M tokens, half of prior GPT-6 Sol pricing. Available via API (
gpt-6.1-sol) and in ChatGPT Work/Codex now. See mapping above. - OpenAI ships Decisions API, routing with GPT-6 Luna (4,494 likes) — replaces if-else app logic with a classify-and-pick call: text or images in, a chosen answer out, 150ms on a Luna variant tuned for the task (vs 1.6s regular). Named use cases: content moderation, request routing, agent next-action selection. Limited to selected API customers now, broader release "very soon." See mapping above.
- PageIndex: tree-search framework hits 98.7% on documents where vector search fails (4,598 likes, Top Repo) — replaces chunk-and-embed RAG with a document tree (like a table of contents) the model reasons through instead of a similarity search. 98.7% on FinanceBench, no vector DB/embeddings/chunking setup, open source and self-hostable, supports MCP.
- Signals list (lower-signal, not individually mapped): Claude Sonnet 5.5 reported matching Opus 5.5 on agentic tasks but burning 7x more tokens to get there; RedAmon's open-source framework that runs full cyberattacks and auto-patches the resulting code; Meta's training trick that doubles Qwen3-8B math accuracy while cutting output length (same model family tracked in [[2026-08-03-innermost-loop-qwen-max-pricing-collapse]]); Blackfrost AI's 180B MoE model (6B active params, GGUF); a locally-run open-source AI companion that plays Minecraft and supports 30+ LLM models.
Zero deep-fetches this issue — the newsletter's own body already carries the specific benchmark/pricing numbers for both Top News items, and the Signals items don't clear the "specific hook + plausible RDCO relevance" bar beyond what's summarized above.
⚠️ Sponsorship
Three distinct paid placements this issue, none overlapping the same-issue editorial content:
- WorkOS — standalone "Presented by WorkOS" block, pitching WorkOS Radar (device/email/network risk-scoring at signup to block free-tier/fake-account abuse before it burns inference credits). Recurring — confirmed pool member since 2026-09-15.
- Sentry — standalone "Presented by Sentry" block, live workshop on instrumenting agent tracing (chatbot, Slack agent, CI PR-review action) to catch bad tool calls and track token spend. Recurring — confirmed pool member since 2026-09-25 (there it ran as a native Signals-list ad; here it's a full standalone block, so placement format rotates even for the same sponsor).
- DigitalOcean & NVIDIA — joint "Presented by" tag inside Signals item 2, promoting the Open Intelligence Summit (SF, Oct 13, with Nous Research and LanceDB) via a partner-link "Apply" CTA. NEW — first appearance of either entity in the tracked rotating pool (now 30+ confirmed distinct sponsors/co-sponsors).
The masthead "In Partnership with" slot is present but unresolved in this issue — plaintext extraction shows the label immediately followed by "Today's Author" with no legible name or link in between, consistent with the majority-unresolved pattern (clean resolutions only on 09-09/QA.tech and 09-25/Voices).
Related
- [[2026-09-23-alphasignal-opus55-gpt6-sol-luna-pricing]]
- [[2026-09-29-mostlymetrics-token-cost-budgeting-frontier-convergence]]
- [[2026-08-04-mostlymetrics-cfo-token-cost-gate]]
- [[feedback_automode_classifier_hard_gate]]
- [[feedback_delegation_model_effort_pairing]]