AI Chatbots Leak Prompts to Ad Trackers: What the IMDEA Study Found — and What Cursor, Copilot, Claude Code, Cline, and Aider Actually Send

privacysecuritycursorgithub-copilotclaude-codeclineaiderlocal-llm

TL;DR: A new academic study from IMDEA Networks — “Prompt like a Butterfly, Sting like a Tracker” — tested nine major conversational AI services and found that 6 of 9 web versions sent conversation-derived data (the conversation URL, title, prompt, or screenshots) to third-party advertising and tracking services, often tied to persistent user identifiers. It hit the Hacker News front page on October 1, 2026 (400+ points, 130 comments). The study covers consumer chatbots, not AI coding tools — but the question it raises applies directly: what do Cursor, GitHub Copilot, Claude Code, Cline, and Aider actually send off your machine? We verified each vendor’s current documented policy on October 4, 2026. Short version: the coding tools look meaningfully better than the consumer chat apps, with one big exception — GitHub Copilot’s Free/Pro/Pro+ tiers train on your prompts and code snippets by default since April 24, 2026, unless you opt out.

ToolCode leaves your machine?Trains on your code by default?Opt-out mechanism
Cursor (Privacy Mode on)Yes — Cursor + model providers (zero-retention agreements)NoPrivacy Mode, available on every plan
Cursor (Privacy Mode off)Yes — may be storedYes, may trainTurn Privacy Mode on
GitHub Copilot Free/Pro/Pro+YesYes, since Apr 24, 2026Account settings opt-out
GitHub Copilot Business/EnterpriseYes — zero retention for IDE completionsNo (contract exemption)Already default
Claude CodeYes — Anthropic APIAccount-level setting (consumer plans)Settings + DISABLE_TELEMETRY=1
ClineOnly to the provider you configureCline itself: neverAnonymous telemetry, opt-out at install
AiderOnly to the provider you configureAider itself: neverAnalytics are opt-in only
Cline/Aider + Ollama (local)NoNoNothing to opt out of

What did the “Prompt like a Butterfly” privacy study actually find?

Six of nine web versions, and three of eight Android versions, of major AI chat services sent either the conversation URL, the conversation title, the user’s prompt, or screenshots to third parties — frequently alongside persistent identifiers that let trackers attribute conversations to a specific user.

The paper — “Prompt like a Butterfly, Sting like a Tracker: A Privacy Analysis of Web and Mobile Conversational AI Agents,” from researchers at IMDEA Networks — ran static and dynamic analysis against nine services on both web and Android: ChatGPT, Claude, Gemini, Grok, Microsoft Copilot, Perplexity, DeepSeek, Mistral’s Le Chat, and Meta AI. The researchers mapped which third-party advertising and tracking services (ATSes) were present, what data flowed to them, and how consent banners, subscription tiers, and access controls changed the exposure.

Two findings stand out beyond the headline number:

  1. Conversation permalinks without access control. Some providers expose shared-conversation URLs publicly, with no authentication — meaning any tracker that sees the URL can read the entire conversation. If you have ever shared a chat link containing code, keys, or internal architecture discussion, that content may be readable by more parties than you intended.
  2. Paying does not reliably buy privacy. The study specifically evaluated whether subscription tiers changed third-party exposure. Tier and consent choices influenced, but did not eliminate, data flows to trackers.

The study was front-page news on Hacker News on October 1, 2026, drawing 400+ points and 130 comments, with the discussion converging on local-first setups as the practical mitigation — the same conclusion our own security series has landed on repeatedly.

Does the study cover Cursor, GitHub Copilot, or Claude Code?

No. The nine services tested are consumer chat products on web and Android. No IDE, editor extension, or terminal coding agent was in scope — and the “Copilot” in the study is Microsoft Copilot, the consumer chatbot, not GitHub Copilot in your editor. Extrapolating the findings one-to-one to coding tools would be wrong, and we’re not going to do it.

The structural difference matters. Consumer chat apps are ad-adjacent web products: they run marketing scripts, analytics, and attribution pixels on the same pages where you type prompts. That is exactly the surface where the study found leaks. An IDE extension or CLI agent talks to an inference API over TLS with no ad-tech in the request path — the tracker-injection vector mostly doesn’t exist there.

But “mostly” is doing real work in that sentence, for two reasons:

  • Coding tools have web surfaces too. Copilot Chat on github.com, the claude.ai web interface, Cursor’s web dashboard and cloud agent pages — these are regular web pages, and the study’s findings about web-page tracker exposure apply to them in a way they don’t apply to your terminal.
  • The real coding-tool privacy questions are retention and training, not ad trackers — and those questions have concrete, documented, changing answers. One of them changed significantly this year.

What does each AI coding tool send, keep, and train on?

Everything below is from vendor documentation current as of October 4, 2026. Policies in this space change — GitHub’s did in April — so treat the verification date as part of the claim.

Cursor: Privacy Mode is the whole ballgame

With Privacy Mode enabled, Cursor’s privacy documentation states your code is never used for training, and Cursor maintains zero-data-retention (ZDR) agreements with its model providers so they don’t store or train on your code either. Privacy Mode is available on every plan, including the free Hobby tier, and is on by default for Teams and Enterprise, where admins can enforce it org-wide.

With Privacy Mode off, Cursor may store codebase data, prompts, editor actions, and code snippets, and may use them to train its models. Two caveats worth knowing even with it on: your code is still transmitted to Cursor and the model providers to generate responses — the guarantee covers retention and training, not transmission — and a small number of models require provider-side retention and fall outside the ZDR agreements (Cursor gates those behind admin approval for privacy-enabled accounts).

If you run Cursor on an individual plan ($20/month Pro, per cursor.com/pricing as verified September 21, 2026) and have never opened the privacy settings, check now. It’s one toggle.

GitHub Copilot: the April 2026 default flip

Since April 24, 2026, interaction data from Copilot Free, Pro, and Pro+ users — inputs, outputs, code snippets, and associated context — is used to train models by default unless you opt out in your GitHub account settings. GitLab’s governance analysis of the change is a useful outside read on why this matters for anyone whose employer hasn’t put them on a Business seat.

Copilot Business ($19/user/month) and Enterprise are exempt under contract terms: for IDE use, GitHub states prompts and suggestions are not retained once the response is returned. Three qualifiers from GitHub’s own documentation: Copilot Chat on github.com (outside the IDE) is retained, unlike IDE completions; pseudonymous engagement data — which suggestions you accepted or dismissed, error events, usage metrics — is kept for two years on all plans; and which company hosts the model varies by the model you pick, which determines whose infrastructure sees your prompt.

If you’re on a personal Copilot plan and never touched the setting since April: your code snippets are training data right now. The opt-out is in Settings → Copilot → Policies.

Claude Code: telemetry and training are two separate switches

Claude Code sends your prompts and code to the Anthropic API to do its job — that part is inherent. Around it, per the official data-usage docs, there are two independent controls that people routinely confuse:

  • Operational telemetry (usage metrics, error reports) goes to Anthropic and third-party logging infrastructure, but never includes your code, prompts, or file paths. Kill it with DISABLE_TELEMETRY=1, error reporting with DISABLE_ERROR_REPORTING=1, or everything non-essential with CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1 in your environment or ~/.claude/settings.json.
  • Model training is governed by an account-level privacy setting on consumer plans (the “help improve Claude” toggle introduced with Anthropic’s 2025 consumer-policy change). Setting the telemetry env vars does not change your training preference, and vice versa. Enterprise and API usage follow separate commercial terms that exclude training by default.

Cline and Aider: the open-source tools are the quietest by default

Cline is an open-source client — your code goes only to whatever provider you configure, which is the entire point of BYOK. Cline’s telemetry documentation is unusually explicit about what its anonymous telemetry never contains: no code or file contents, no file paths or names, no conversation content, no API keys. You’re prompted about it at install.

Aider goes one step further: analytics are opt-in only, and Aider’s analytics docs state it never collects your code, chat messages, or keys — just model/feature usage tied to a random UUID when you agree. Among the five tools here, Aider has the most conservative default posture, full stop.

Does a local model actually give you zero egress?

For the inference path, yes. Point Cline (or Aider, or Continue.dev) at an Ollama server on localhost and your prompts and code never cross a network boundary:

# Terminal 1: serve locally (binds 127.0.0.1:11434 by default)
ollama serve

# Verify it's answering locally, not via any cloud relay
curl -s http://localhost:11434/api/version
# → {"version":"0.12.x"}

In Cline: API Provider → “Ollama”, Base URL http://localhost:11434, pick your pulled model. Our privacy-first Cline + local LLM setup guide walks through the full configuration, and the LM Studio privacy guide covers the GUI-server route.

Here’s the problem we hit when we actually watched the traffic, though: a local model does not mean a silent machine. Running Cline against localhost Ollama inside VS Code, the editor itself still phones home — VS Code sends crash, error, and usage telemetry to Microsoft independently of anything your AI extension does. The fix is one setting: "telemetry.telemetryLevel": "off" in settings.json, which per VS Code’s telemetry documentation also covers first-party extensions — but third-party extensions are only asked to respect it, so audit what you’ve installed (our Open VSX evil-twin extensions guide covers why that audit matters for security, not just privacy). Cline’s own telemetry is a separate opt-out, as above.

The honest cost line, because a how-to without one is useless: local coding models good enough for daily agentic work need VRAM. A used RTX 3060 12GB ($260–$296 street, September 2026) runs small coder models acceptably for autocomplete and single-file edits; comfortable 30B-class agentic coding wants a used RTX 3090 24GB, which the DRAM-crisis market has pushed to $1,150–$1,350 as of September 2026. Whether that’s worth it versus a $10–$20/month subscription with the opt-outs configured is a budget decision, not a privacy decision — runaihome.com’s VRAM-to-model guide has the hardware math, and aifoss.dev’s Ollama review covers the server side.

Honest take: what should you actually change this week?

Change nothing if you’re on Copilot Business/Enterprise, Cursor with Privacy Mode enforced, or Claude Code under commercial terms — your contractual posture is already stronger than anything in the consumer chat apps the IMDEA study tested.

Spend five minutes on toggles if you’re on a personal plan — this is most readers. GitHub account → Copilot → opt out of model training (the April 24, 2026 default affects you). Cursor → enable Privacy Mode. Claude Code → check the account-level training setting, and set CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1 if you want the quiet version. VS Code → telemetry.telemetryLevel: "off". None of this costs features.

Go local BYOK only if your threat model genuinely requires zero egress — client IP under NDA, regulated code, or you simply refuse to transmit source to any third party. Cline or Aider + Ollama is the setup, the hardware floor is roughly $260 (used RTX 3060) and the comfortable tier ~$1,200 (used RTX 3090), and you still need to silence the editor around the model.

One more thing the study should change: stop sharing conversation permalinks that contain code. Public, unauthenticated chat-share URLs were one of its ugliest findings, and that one applies to developers directly — a shared ChatGPT or Claude link with your auth middleware in it is a disclosure, not a collaboration feature. This is the tenth entry in our security series; the MemGhost memory-poisoning analysis, the Shai-Hulud session-hijack worm write-up, and the VibeDefend governance review cover the attack side of the same trust boundary this study measures from the data side.

FAQ

Did the IMDEA study test Cursor or GitHub Copilot?

No. It tested nine consumer chat services (ChatGPT, Claude, Gemini, Grok, Microsoft Copilot, Perplexity, DeepSeek, Le Chat, Meta AI) on web and Android. The “Copilot” tested is Microsoft’s consumer chatbot, not GitHub Copilot. The coding-tool scorecard in this article comes from each vendor’s own current documentation, verified October 4, 2026.

Is GitHub Copilot training on my code right now?

If you’re on Copilot Free, Pro, or Pro+ and haven’t opted out since April 24, 2026: yes, prompts, suggestions, and code-snippet context are used for training by default. Business and Enterprise plans are contractually exempt. The opt-out is in your GitHub account’s Copilot settings.

Does DISABLE_TELEMETRY=1 stop Anthropic training on my Claude Code sessions?

No — that variable only stops operational usage metrics, which never contained your code anyway. Training on consumer plans is controlled by the account-level privacy setting in your Claude account. They are independent switches; set both if you want both.

Is Cline + Ollama really zero egress?

The inference path is: prompts and code go to localhost:11434 and stop there. The remaining egress is around the model, not through it — VS Code telemetry (set telemetry.telemetryLevel to off), extension update checks, and any other extensions you’ve installed. Cline’s own anonymous telemetry contains no code or conversation content and can be declined at install.

Sources

Last verified October 4, 2026. Privacy policies and training defaults change frequently — GitHub’s did this April. Re-check the official pages above before relying on any default described here.

Was this article helpful?

Know which coding tool is worth paying for

Hands-on comparisons of AI coding assistants and what each one costs to run — including the local-model path. Sent only when something changes. Unsubscribe anytime.