Kwik Trip AI architecture — consolidated engagement context
Compressed from founder-relayed material 2026-08-04 so the context survives beyond the session. Working artifacts (client-facing skeleton + running notes) live OUTSIDE the vault at ~/.claude/state/kwiktrip-architecture-skeleton.md and ~/.claude/state/kwiktrip-architecture-outline.md.
The ask
Adam Gruenke (phData account lead): two asks from Lalan — (1) backlog visibility/burn-down help on a 750-item Jira backlog (some 18 months old, DE + analytics work); (2) turn KT's AI architecture from accidental to intentional. Founder owns #2: initial artifact due EOW Thu 8/7, Friday review call with the program team proposed. Context: the broader phData engagement is an 8-10 week AI enablement evaluation whose expected output is exactly "AI architecture diagram + roadmap" — the founder's doc is the engagement's central artifact, not a side memo. Teams already interviewed by phData (May): Governance, Power Platform, Snowflake/Databricks, ML, Reporting, Engineering.
Platform decisions (settled)
- Snowflake sunset DECIDED week of 2026-07-27. Consolidating on Databricks. Cascades: (a) Sigma exiting (sat on Snowflake); (b) Alation catalog (just launched, connected to Snowflake + PBI) must re-point; (c) in-flight MS SQL→Snowflake SQL translation work now targets a dying platform.
- PBI + Databricks Genie = stated direction (heavy Microsoft shop; PBI P1 capacity). Fabric licensed but unadopted (data-duplication concerns; Microsoft pushing enablement).
- Heavy Databricks users but AI features unused; Genie usage developing. Unity Catalog depth UNVERIFIED (load-bearing for the target state).
Platform ownership map (silos are organizational, drawn at director level)
| Platform | Owner | Org chain |
|---|---|---|
| Databricks | Lalan Kanade, Director of IT Data Services (21) | CFO Dave Wagner → VP IT Bruce Bingham → Sr Dir Chris Halverson → Lalan |
| Power BI | assumed Lalan's data analysts (ex-Sigma devs) — UNCONFIRMED | under Lalan |
| Reporting team | Bryce Reynolds (Data Analyst Mgr) → Jason Dirnbauer et al. | under Lalan |
| UiPath | Rohan Crain, IT BA Manager (4: Bennett Clark, Brandon Burt, Ian Marquez, Kristen Barney) | Bruce Bingham → Brennan Murphy (Dir IT PM) → Leslie Schweitzer → Rohan |
| Power Automate + Copilot Studio | Tyler (last name unknown — harvest Wed), IC Power Platform Admin (no delivery team — thin-ownership surface) | see [[2026-07-17-kwik-trip-engagement-notes-stakeholders-and-priorities]] |
| Both chains meet at Bruce Bingham (VP IT) → natural table for architecture-level decisions. "Allay" from 7/17 notes = Lalan Kanade (resolved; vault note corrected). |
AI governance (May 15 deep-dive, compressed)
- Committee EXISTS (~9 months, meets 1-2wks): Kerri Bednarchuk (data governance, lead voice), Jean Seely (PM), Tyler, Chris Halverson + Demetro (relay grouped both under "security" — but the org chart has Halverson as Sr Dir IT above Lalan; dual role or relay grouping artifact, UNCONFIRMED), Dennis Houston + Brent Reinhart (dev), David Stratton, Viktor, Justin. Founder 8/4: "formed, but their involvement/process needs more forming — they are looking for help there."
- Self-identified gaps: no scoring/ROI model, no charter, no metrics, slow turnaround, perceived as friction, vacant AI program manager role. Intake (Teams form) paused, ServiceNow transition decided-not-started. They floated a "do not pursue" tier themselves.
- Alation: business catalog live, Kerri's team curates, rehiring the key support role (second under-resourced team in the capacity math).
- Copilot deliberately restricted from broad SharePoint search — "data accuracy > data availability" is their own governing phrase.
- Andrew Kim (phData) seeded control-plane/observability/drift language in May — Horizon-3 continuity.
Reporting team (May 14, compressed)
- PBI (P1) + SSRS + Sigma. Certification tags UNUSED (quick win). Jira intake via BA stories, manual updates (some Rovo).
- Jason Dirnbauer (Data Analyst 2, under Bryce/Lalan) designing a multi-agent pipeline concept: Jira intake → requirements/change control → documentation → SQL build → validation → deployment. Uses Copilot, Cursor (personal), Power Automate; Copilot Studio primary; SSRS-XML-through-Copilot workaround; MS SQL→Snowflake translation loops "9/10 success." In-house first customer for the paved road.
Engineering team (May 15, compressed)
- DE team grew ~5→20. Owns pipelines, ML support, data delivery across Snowflake (DW) + Databricks (compute/ML).
- Tools: Cursor = majority usage, VS Code+Copilot limited, Databricks Assistant frequent, Cortex ad hoc, Genie early. Skills/rules/MCPs built individually — no central repo, no agent SDLC (explicit decision: not investing in agent SDLC yet; human-in-the-loop required for all AI outputs).
- GitHub bug bot runs daily, assigns findings to last committer (their most mature agent).
- Docs: Cursor+Confluence, Atlassian Rovo.
- Their own words on the keystone: "weak semantic layer today — metadata/definitions fragmented across Sigma, Snowflake, dbt; no unified enterprise context." (dbt is in the stack.) Pain points: standardization, integration, governance, adoption.
- phData recs already delivered: adoption > tool enablement; Toolkit & Forge shared skills.
- Side use case raised: detecting AI-assisted candidate responses in recruiting tests (Vijay + David follow-up).
ML team (May 15, compressed)
- Flagship production ML: ~500-store food-demand forecasting → automated kitchen production schedules. Closed-loop, NO store-level human review. Prophet + emerging XGBoost, Python in Databricks, Airflow orchestration; data originates in Snowflake, dbt feature engineering; stack spans Snowflake/Databricks/Airflow/dbt/Spark/Ray. People: Jim Havlicek, Jorden Fread, Ryan DeFreitas; Viktor Drozd = phData report recipient.
- Pains: no standardized model-validation metrics · testing gaps · explainability/root-cause of forecast errors · experiment diagnostics shallow · QA env not aligned with prod (env drift) · multi-repo/multi-platform AI context limits.
- Tool preference: embedded platform-native AI (Cortex, Genie, Databricks Assistant); Cursor for broader code changes; default/auto model selection.
- Implication: Snowflake sunset's third domino — the production ML feature pipeline migrates, highest-stakes item. HITL "all AI outputs" decision is genAI-scoped (classical ML already closed-loop) — doc states scoping explicitly.
Cursor & Claude usage (May 15, compressed)
- Cursor: 100+ licensed users, ~150 installs, IT-wide (engineers, QA, PMs, BAs, network, cloud). Claude not directly deployed — Anthropic models consumed via Cursor. Users cycle Cursor/Copilot/Rider on limits.
- Productivity self-reports: 4h→~1h+review · 30-50% task gains · 1-2 wks faster projects. Cost spikes + wrong-mode misuse observed; usage/spend monitored manually via logs.
- Named needs: centralized guidance/standards · persona-based training (David Stratton driving) · control plane (agent ops + FinOps + observability) · eval framework ("deterministic + rubric + SME") · enterprise AI strategy alignment. Moonshots: retail-ops agent, Teams alerts for agent issues, clean data foundation for MCP/agents.
- Founder also shared phData's "Dual Exponential" advisory diagram (efficiency vs reinvention curve, harvest→redirect→compound motion, "5.7 hrs saved / 1.7 redirected" stat) — candidate framing for the Argument section's "what intentional buys."
Agent portfolio (June 18 discovery calls, compressed — founder-pasted 8/4)
- Bread & Bun Mix Assistant (in delivery, the first agent): bakery operator (Kevin Ranzenberger, sponsor Joe Blatz) built a Copilot chat classifying flour farinograph specs (MTI + Stability → red/yellow/green/purple zones) into exact mix instructions + amp targets. phData Phase 1 = preserve-and-structure chat → agent (Taylon); Phases 2-4 = structured data (COA, silo/MES/OnBase, ~3yr production history) → formal predictive modeling in Databricks → closed-loop. The routing rule walked end-to-end by their first real agent; graduation trigger = precision. Chat→agent extraction prompt saved for reuse (Jean) = first paved-road on-ramp artifact. KPS (store forecasting) cited as the Phase-3 precedent.
- Risk agents (specced June 18, NOT started, priority "nice to have"): 4 use cases — body-shop/repair-estimate validation (MVP, rule-based Copilot; fuel→mechanical, car-wash→exterior), early claim risk/cost estimation (PII-heavy, elevated tier), settlement/reserve validation (split as separate case), snow slip&fall liability transfer (Salesforce CLM contracts, legal risk, human-validation mandatory). People: Austin Fenzl (claims data), Katrina Jessie, Tom Colbert, Leslie Schweitzer; third-party vendor Klear AI (Sohan Balcharan) in the loop for data feeds/PII redaction.
- Systems learned: Inform (claims RMIS) has NO direct connector — MCP server or doc-export discussed; Salesforce CLM = easy Copilot integration; recurring pattern = data access, not AI capability, is the blocker.
- phData already teaching the platform-fit split (Taylon, June): Copilot Studio = simple rule-based · AI Foundry/Databricks = advanced (image/vision, formal prediction). Routing rule = codifying existing practice.
- Governance risk-tiers already informally applied (PII/legal cases → "higher risk tier," governance + legal review, HITL mandatory).
AI Initiatives Jira board (48 tickets, founder access 8/4 — distinct from the 750 delivery backlog)
- Org-wide idea intake, stale (overdue since Jan/Apr). Double-clicks: AI-12 Relex / AI-11 KPS / AI-23 demand-planning ALL EMPTY (no context) — intake-quality evidence; the Relex buy-vs-build question stays open for Wed.
- AI-22 Eagle Eye: David (Marketing) → Brennan email. Personalized loyalty challenges — vendor platform uses AI on purchase data to push scaled challenge portfolios ("100s running"); KT feeds data + iframe in app, avoids AMS integration; possible vendor funding. Status: parked until marketing team rehired. High interest.
- AI-6 video safety: BUY-not-build expectation (ties into existing surveillance; spills/lifting/trip hazards). Dept 669, contacts Chris Lenser/David Stratton/Cassie Chambers, "just an idea," no due date. (Corrects my earlier "routing miss" read — it's a vendor candidate despite the Copilot Agent label.)
- KEY ORG FACT: Brennan Murphy "started his AI group" (per AI-6 ticket) — Brennan (Dir IT PM, Rohan's chain) founded/runs the AI group; likely = the "AI Enablement Team" in calendar invites. Doc audience/champion alongside Bruce's table.
- Doc upgrade from this batch: routing rule gains a 4th BUY/EMBED lane (vendor-embedded AI: Eagle Eye, video safety, Relex?, Klear AI) — KT's build = governed data feeds + integration surface; gate = committee's existing vendor-evaluation policy; semantic layer's consumer list extends to vendor data contracts.
Doc architecture (founder-ratified through v1.1; v1.3 adds ML flagship in §3, HITL scoping, ML-migration Horizon-1 item, Cursor/productivity texture in 5 T's; v1.4 adds Bread & Bun case study in §3, business-systems inventory incl. Inform/SF-CLM/Klear AI, risk-tier axis in §5, capability+precision graduation trigger)
7 sections: Argument (Snowflake sunset as cascade-proof hook) · Current state (overlap matrix = lead visual: patterns × platforms × owning team) · Use-case patterns (paved roads, routing rule) · Target state (Databricks gravity; semantic layer keystone — Metric Views feeding BOTH PBI and Genie; Fabric fork adjudicated toward Databricks-semantic→PBI-publish; UC spine + Alation role split; anti-bottleneck self-service tiers) · Operating model (EQUIP the existing governance committee, altitude split: use cases → committee, architecture forks → Bruce's table) · Sequencing (3 horizons; Sigma rebuild once-on-semantic-layer; retarget SQL work to Databricks during migration) · Appendix: 5 T's dictionary (Time/Talent/Treasure/Tasks/Tools). Capacity math (design constraint): Lalan's 21 carry migration + Sigma→PBI rebuild + Genie enablement + AI intake; Kerri's team carries Alation re-point while rehiring.
Open items (Wed 8/5 onsite harvest)
- UC adoption depth (real catalogs/permissions vs partial) — determines Horizon 1
- PBI tenant/workspace ownership confirm
- Snowflake→Databricks migration timeline
- Copilot Studio + UiPath data-source connections
- Backlog export → bucket 750 by theme same-day
- Predictive-ML pattern examples; first-90-days list
- Possible additional mine: phData's Power Platform + ML team interview notes (May round)
Related
- [[2026-07-17-kwik-trip-engagement-notes-stakeholders-and-priorities]] — stakeholder map (Tyler, Lalan resolution)
- [[2026-07-31-kwiktrip-onsite-strategic-prep]] — travel/onsite logistics
- [[2026-07-31-sprocket-document-pipeline-tool-selection]] — Sprocket thread (Databricks/UiPath/Copilot Studio overlap precedent)
- [[2026-07-13-kwik-trip-deskless-agent-distribution]] — deskless distribution analysis