Skip to content
Why GEO Why us Pricing
Get started free
Blog/Case Studies
Case Studies

Case study: Stack Overflow — being in the model isn't being visible

Stack Overflow is the clearest documented case of a brand feeding the models and losing the audience: monthly questions fell from a ~200,000 peak to levels last seen at its 2009 launch, even as its answers kept working inside every coding assistant. It's the sharpest proof on record that being in the model and being visible in its answers are entirely different things.

Portrait of Maya Lindqvist Maya Lindqvist · Head of Research Aug 9, 2026 9 min read
STACK OVERFLOW — QUESTIONS ASKED PER MONTH~200k · 2014 peak2009 levels · 2025Shape traced from public Stack Exchange data; see sources
Key takeaways
  • Monthly questions on Stack Overflow collapsed from a peak of roughly 200,000 to levels last seen in 2009, the year it launched — with the steep leg starting right after ChatGPT's November 2022 release.1,2
  • The audience didn't leave programming Q&A — 84% of developers now use or plan to use AI tools, per Stack Overflow's own 2025 survey.3 The need moved into the assistant.
  • In May 2024 Stack Overflow licensed its corpus to OpenAI — monetising the archive without restoring the visits.4 Being training data pays once; being the recommended answer pays continuously.
  • For brands, the operative distinction is ingredient vs destination: engines consuming your content is not the same as engines sending you buyers — and only one of them is worth optimising for.

Stack Overflow lost its audience to models trained substantially on its own content — and nothing about its SEO broke. Its pages still rank. The knowledge is still consulted millions of times a day, inside ChatGPT, Copilot and Claude, without the visit. Where Chegg is the case study in losing paying customers to the AI answer, Stack Overflow is the stranger and more instructive one: proof that a brand can power the answers and still vanish from them.

The numbers, sourced

Stack Overflow's activity data is public, which makes this collapse unusually easy to verify:

💡

Note what the metric is: questions asked, not just visits. The substitution reached the contribution loop itself — the site's future content supply — which is what makes this decline structural rather than cyclical.

Meanwhile the demand didn't shrink; it relocated. Stack Overflow's own 2025 Developer Survey found 84% of developers use or plan to use AI tools in their workflow, up from 76% the year before — and, in a detail the company itself highlighted, only 29% trust the output.3 Developers still ask the same questions. They ask them somewhere that answers instantly and never marks a question as duplicate.

The OpenAI deal: monetising the archive, not the visibility

In May 2024, Stack Overflow and OpenAI announced an API partnership: OpenAI pays for access to the question-and-answer corpus through OverflowAPI and surfaces Stack Overflow content in ChatGPT.4

This is a rational response — if the model is going to consume your content anyway, get paid. But it's worth being precise about what the deal is: it converts the archive into a licensing asset. It does not restore the visits, the question flow, or the ad and talent-product funnels those visits fed. One asset was licensed; the other — being the destination — was already gone.

Ingredient vs destination: the GEO lesson

Stack Overflow is the cleanest demonstration of a distinction every content-producing brand should internalise:

IngredientDestination
Your content is…absorbed into training data or retrieval, anonymouslynamed, cited and recommended in the answer
The buyer sees…the knowledge, with no idea it came from youyour brand as the thing to choose or visit
It pays…once, if you can negotiate a licence — nothing otherwisecontinuously, in buyers who act on the recommendation
You can measure it by…essentially nothing — training data is a black boxmention, citation and recommendation rates per prompt

"We're all over the training data" is not a visibility strategy — Stack Overflow was the training data for its category and still lost the audience. Visibility that pays is the destination kind: when a buyer asks which tool should I use, the engine names you, and independent measurement shows engines do send destination traffic to brands they recommend — that's the pattern in the Vercel case, the winning mirror image of this story.

The practical corollary: aim your effort at the pages engines cite as sources — the comparison, spec and answer pages where being named is the point — rather than donating undifferentiated knowledge to the commons. Which sources engines actually cite for your prompts is measurable: the Answers view reads them off every tracked answer, and Pew's data shows the citation economy is real but concentrated — Wikipedia, YouTube and Reddit topped the cited-domain list in Google's AI summaries.5 We've written about what that concentration means for brands.

What to watch on your own site

  1. The ratio of consumption to attribution. If AI crawlers are fetching your content (check who's allowed in with the AI crawler access tool) but engines never cite your domain in answers, you're an ingredient — the Stack Overflow position, minus the licensing leverage.
  2. Prompt-level presence. On the questions your buyers actually ask, are you named? Recommended? Or is the answer complete without you? Track it engine by engine — here's how we measure it — because the engines differ.
  3. The trend, not the snapshot. Stack Overflow's decline was visible in its public data for anyone who charted it quarterly. Your equivalent chart exists too; the only question is whether you're the one reading it.

Frequently asked questions

It monetised the archive — OpenAI pays for API access to the corpus — but it didn't restore visits or question volume. Licensing your content to the model and being visible in its answers are different assets with different values, and the deal only priced the first one.

Same mechanism, different funnel. Chegg lost paying subscribers when AI answered homework questions; Stack Overflow lost the contribution loop — the questions that generated its future content. Both are substitution cases: the AI answer met the need the visit used to meet.

Not whether your content is "in" the training data — you can't control or meaningfully measure that. Measure what engines actually say: whether you're named, cited and recommended on the prompts your buyers ask, and which sources the answers draw on. That's the visibility that sends buyers, and it's measurable.

Sources

  1. Slashdot — "Stack Overflow Went From 200,000 Monthly Questions To Nearly Zero", January 5, 2026, summarising public Stack Exchange activity data.
  2. Gergely Orosz, The Pragmatic Engineer — "Stack Overflow is almost dead", May 2025 — charts monthly questions back at 2009 levels using the public data explorer.
  3. Stack Overflow — 2025 Developer Survey, AI section and press release: 84% of developers use or plan to use AI tools; trust in AI accuracy at an all-time low (29%).
  4. OpenAI — "API Partnership with Stack Overflow", May 6, 2024.
  5. Pew Research Center — "Google users are less likely to click on links when an AI summary appears in the results", July 22, 2025 — including the most-cited-source breakdown (Wikipedia, YouTube, Reddit).
Share
Portrait of Maya Lindqvist
Maya Lindqvist

Head of Research at MentionBeat. Maya leads the measurement methodology behind MentionBeat's visibility metrics — prompt-suite design, sampling, and confidence intervals — and writes about how generative engines choose what to say.

Are you an ingredient or a destination?

MentionBeat measures whether engines name, cite and recommend you on real buyer prompts — and which sources their answers draw on.

Get your free visibility report
No credit card. Results in about a minute.