06-reference/transcripts

indy dev dan is anthropic stealing your data transcript

2026-07-27

Transcript: "Is Anthropic STEALING Your Data? (While You PAY FOR IT)" — IndyDevDan

Engineers, small business owners, and enterprise leaders, it's time we sit down, look at the elephant in the room, and answer the question head-on. Is Anthropic stealing your data while you pay for it? This video and this channel, it's not hype or fear-based. If we can't answer the question without supporting ground truth data, we don't answer at all. But, there is sufficient evidence to support the fact that Anthropic is competing with its customers, and I'm not the only one asking this question. Satya Nadella, CEO of Microsoft, says this, "You essentially pay for intelligence twice. Once with money and again with something even more valuable, the proprietary knowledge you must reveal to make that intelligence useful. The better you want the model to perform, the more of that knowledge you have to feed it." This is a brutally true take on AI and on agents. And that's not all. We have Alex Karp, CEO of Palantir, saying this,

[00:01:01] "What the technical customers want is control over their compute, their models, their data stack, their alpha." Keyword, alpha. "They want to know that the means of production is not being transferred to someone else. Who owns the data? Are the prompts secure? Is this being transferred to you?" Microsoft even went as far to ban Fable 5 over the 30-day ZDR data retention policy. I don't agree with both of their takes 100%, but their views and points are valid. I've been engineering for over 15 years now, and I've become hyper aware of companies that use my data and use your data to fuel their business. Language models and agents are far worse than Google Search. It's not just the query we're handing over to these companies. It's, as Satya and Alex Karp mentioned directly, it's the intelligence of our business and therefore our livelihood. In this video, we answer the question directly, "Is Anthropic stealing your data while you pay for it."

[00:02:03] Let's visualize the problem properly. Your business hands over your IP through prompts, agents, traces, workflows, and we pay twice here. As Satya said, you're paying in cash and you're paying the IP you hand over. This is the problem. If this is happening, it's only a matter of time until every model lab takes over every industry. So here's what we're going to cover: answer whether Anthropic is stealing your data, the aggregate bucket problem, the ground truth facts, commodity agents versus IP agents, and the solution — open models being a key piece.

[00:03:00] So is Anthropic stealing your data? No, they are not. But they use it anonymized in aggregate. Think about a data tumbler — all of our prompts go into a bucket, they anonymize it, and now they have aggregate trends. When you prompt in Claude Code or Codex, you are renting the model.

[00:04:00] It's still a map of the market — coding, design, legal. There is a chain of evidence of Anthropic competing with its customers: Cursor, then Claude Code; Figma MCP usage, then Claude Design; security usage, then Claude Security; life science, then Claude Life Science. One is a coincidence, four is a pattern.

[00:05:01] This is how it works: they see the usage, they learn the trends, they enter the vertical, then they cut access. They're not cutting access much anymore (some Cursor and Windsurf access was cut historically), but the principle stands — you and I do not own the model. This is dependency risk, keyman risk, key technology risk.

[00:06:00] Fable was recently banned by the government (referencing Microsoft's ban of Anthropic's Fable model over the 30-day ZDR retention policy). This is in the terms of service — you and I agree to this. Anthropic's Clio system does "aggregated privacy-preserving analysis of data to gain insights into the real-world impacts of AI systems... while maintaining user privacy."

[00:07:01] Anthropic's economic indexes come from Clio. In Claude.ai settings > privacy, there's a toggle: "Anthropic may conduct aggregated anonymized analysis of data to understand how people use Claude" — recommend opting out of "help improve our AI models." Props to Anthropic: they don't sell data to third parties and delete data on request.

[00:08:01] There's a 30-day required retention policy for the Fable model due to cybersecurity harms. Not all customer classes are equal — consumers vs. commercial (teams, enterprises, API users). Free-plan users: "you are the product." Pro and Max data still feeds Clio's aggregate reports.

[00:09:01] Not all customers are protected equally — commercial is much better defended than consumer. The only thing keeping AI labs honest on privacy is incentive, not goodwill: violate a term of service once and there's a mass enterprise exodus.

[00:10:00] Four claims tested: (1) they see aggregate patterns — true; (2) they train on your code/prompts — false per ToS; (3) they own your outputs — false, ToS says you own them; (4) they compete with your product — true, if your domain is scaling and profitable.

[00:11:00] Verdict: they are not stealing your data, but they are competing with their customers — could be you, given the right numbers, growth, success. What a single engineer, team, or agent-first org can accomplish today is exponentially higher than a few years ago.

[00:12:00] This is an unprecedented problem engineers must face. The channel's approach: don't fear-monger, offer the solution. When you hit enter, all that agentic engineering work gets sent to the lab, stuck in a database.

[00:13:01] Two types of agentic coding: commodity agentic coding (prototypes, CRUD, boilerplate, glue — doesn't matter if labs see it) vs. IP agents (business know-how, prompts, traces, domain logic, hard-earned evals, user data insights — scarce, asymmetric, compounding).

[00:14:00] The core test: "If a competitor could read my full agent trace, would it matter?" Most engineering work is commodity work. IP-dense work is the sub-fraction of a percent distribution the models have never seen in training data. Privacy is a stack: account, contract, feature, model, routing, and the agentic engineering around it.

[00:15:01] Most engineers aren't doing IP work, and that's fine. For everyone else, climb the "AI sovereignty ladder." Recommend individuals/small businesses/enterprises reach tier three, and tier four if possible: hybrid private — open-weights model on a rented GPU stack you control, owning your own endpoints. Own the control plane at minimum (VM + LLM gateway).

[00:16:02] Model cloud (AWS, GCP Vertex, Microsoft Foundry) offers real protection — Anthropic steps away once mounted on a cloud provider's infra, so you go through the cloud provider instead of the lab directly. Below that, the commercial API is the lowest protection tier still better than subscription — Anthropic explicitly states it only samples aggregate data from the API.

[00:17:00] For individuals: stay at subscription level or move to commercial API. If an individual is above ~$1M revenue and doing real agentic (non-commodity) work, move up the tier stack. For small/medium businesses (2-100 people): move to tier two or three — model cloud means your data isn't sampled at all.

[00:18:02] For enterprises: own the control plane and set up open-weight models on rented GPU infrastructure. "If you're doing real agentic engineering work that's scaling, you got to own the model... it could be snapped away from you" at any lower tier (government bans, ToS changes).

[00:19:01] Tier four (open weights on rented/owned GPUs) is the real solution as businesses scale — anyone can prompt a model, very few can build software factories and developer workflows; that's where true IP gets built. Full on-prem GPU ownership is the top tier but "basically impossible" for most — huge capital requirement.

[00:20:01] Open-weight Chinese models (Kimi K2, GLM, MiniMax, Qwen) are commoditizing closed-source US labs, but the speaker is skeptical of using overseas API providers directly — ToS there can't be reliably enforced, and there's a real US-China AI competition dynamic. Self-hosting wins on secrecy and trace ownership, not automatically on price.

[00:21:01] Best/mid/worst case scenarios: best — labs keep behaving as documented; mid (speaker's bet) — labs stay disciplined platforms focused on models/GPU/chips/energy rather than climbing the application stack; worst (speculative, not believed) — anonymization becomes a legal loophole where "anonymized" data effectively becomes lab IP once scrubbed of ownership markers. The defense is the same regardless: own the traces, own the evals, keep a second model path for IP-essential work, rent GPUs, own the models.

[00:22:00] Final answer: Anthropic is not stealing your data — the public record and incentive structure say that's false. But your usage is shaping what they build next. Verticals already seen: Claude Code, Claude Design, Claude Life Science, Claude Security — likely more coming (Claude Legal, Claude Animation, etc.). Same verdict applies to OpenAI and any model-lab platform.

[00:23:00] The sweet spot: rent GPUs from a cloud service, run an open-weight model on top, own all the traces — the only data sent out is GPU usage. Set up your own router/API to switch between models while owning the learning loop.

[00:24:00] Closing: "Is Anthropic stealing your data? No, but they are using it to make plays vertically in the domains that are making the most cash." This is the platform play generally, not unique to Anthropic — "we would do the same thing in their position." Sign-off: weekly Monday uploads for engineers building in the age of agents.