06-reference

technically open weight models part2

2026-09-01·reference·source: Technically·by Will Raphaelson
open-weight-modelsmodel-vs-harnessquantizationlocal-inferenceharness-engineering

Why this is in the vault

Part 2 of Technically's open-weight series — a practical dial-by-dial guide for picking a model (size, quantization, context, reasoning, tool-calling) and a harness — that turns the "model choice is commoditizing" thesis from the part-1 note into an operational checklist RDCO can actually apply.

The core argument

Raphaelson picks up from part 1 ([[2026-08-18-technically-open-weight-models]]) with a practitioner's framework for choosing an open-weight model + harness combo, organized around five tunable dials that jointly determine a model's fit for a task:

Worked example: a fridge tracker that texts you about spoiling food needs none of the heavy dials — small model, 4-bit, small context, no reasoning, no tools. A significant software-engineering project or a math paper needs the opposite setting on every dial. He anchors the two extremes with named models: Qwen 3 1.7B for the small/fast end, Kimi K3 (2.8 trillion parameters) for the large end — the same Kimi K3 covered in the production-engineering piece already in the vault.

On harness choice: harnesses split into general-purpose (chat/coding) and an emerging class of domain-specific harnesses that bake in tools, prompts, and guardrails for one job — Harvey (legal), Figma AI (design), Abridge (clinical notes), Sierra (customer service), Cursor's Bugbot/Greptile (code review), Clay (GTM/lead enrichment), Glean (enterprise search with citations). For open-weight experimentation specifically, he recommends free harnesses built to plug into open models regardless of where they're hosted: LM Studio and Jan.

Note on completeness: the Gmail plaintext body cuts off mid-piece at "Small Model Usage Example..." — the walkthrough itself (presumably a concrete LM Studio/Jan session running Qwen 3 1.7B) is on the web post but not captured here. The model-picking and harness-picking frameworks above are the complete, load-bearing argument as delivered; the cut section is a worked demo of the same framework, not new conceptual content.

Mapping against Ray Data Co

This is the operational sequel to the harness-engineering thesis already anchored in the vault ([[project_l5_north_star_strategic_direction]], [[2026-07-21-technically-harness-engineering]]): where part 1 established that model quality is commoditizing and the harness is the moat, this piece gives RDCO the actual decision procedure for the model side of that split, which matters directly for two live threads. First, client-facing model-choice conversations at phData/OI: the five-dial framework (parameters/quantization/context/reasoning/tools) is a clean, non-hand-wavy way to walk a client through "why not just use the biggest model for everything" — RDCO's own Claude-default posture is itself a dial-setting decision (large context, tool-calling on, reasoning on for agentic work) that this framework makes explicit and defensible. Second, cost/capability tradeoff decisions for RDCO's own tooling — the vendor-choice gap flagged in the inference-providers note ([[2026-04-16-technically-inference-providers]]: "worth a future audit when API spend becomes a real budget line") now has a concrete first move — a small quantized open-weight model (Qwen 3 1.7B class) is the right dial-setting for RDCO's high-volume, low-complexity slices (newsletter triage classification, simple extraction) versus the current frontier-lab-default-for-everything posture, if/when that audit happens.

Related