Google AI Hallucination Problem: Why Search Results Lie and How to Fix It

Skep spots a hallucinated Google AI overview — confident wrong answers at the top of search results

TL;DR: Google’s AI overviews hallucinate on 15 to 40% of long-tail factual queries, presenting confident but false information at the top of the results page. The problem hits hardest for obscure technical, historical, or local facts. The fix doesn’t come from Google itself, but from what you do in the next ten minutes.

Google’s AI answers are wrong more often than you think — and it’s not random

Open a Google search for an uncommon historical event, a niche Python library edge case, or a local regulation in a mid-sized city. The AI overview at the top may read as a clean summary, but the odds that it contains a material error are not low. Public hallucination benchmarks from 2026 show that frontier models, the class of systems powering these overviews, hallucinate on approximately 15 to 40% of long-tail factual queries. Head-of-distribution facts — capital cities, major dates — see only 1 to 3% errors, but the internet’s real value lives in the long tail. When you search for something specific because you need to get it right, the AI answer is disproportionately likely to be wrong.

The user backlash captured in the Reddit thread “Google has officially gone insane” isn’t just frustration with change. It’s a direct consequence of confident errors appearing where authoritative information used to be. Users describe adding “reddit” to queries because they learned that crowdsourced human discussion, however noisy, outperforms a model that doesn’t know what it doesn’t know. The same thread surfaces a migration toward DuckDuckGo and other non-AI search engines, not out of ideological preference but because getting the right answer matters.

The problem isn’t that AI sometimes makes mistakes. It’s that Google surfaces AI-generated text as a definitive answer, above the organic links, with no visual distinction that says “this might be fabricated.” A user who doesn’t scroll past the overview gets a plausible-sounding falsehood and no signal to doubt it.

Why the real cause is benchmark incentives, not just model flaws

The standard explanation is that models hallucinate because training data is imperfect or reasoning is incomplete. That’s true but superficial. The deeper structural problem is that the benchmarks used to evaluate these systems measure what vendors choose to measure, not real-world reliability in user workflows.

Consider the Mirage effect documented in multimodal AI evaluation. Researchers showed that frontier models generate detailed image descriptions and clinical findings even when no image is provided, a phenomenon they call mirage reasoning. The models fabricate plausible perceptual narratives by exploiting textual cues and hidden structure in benchmarks. The same mechanism operates in search: given a query string, Google’s AI constructs a coherent-sounding answer that matches the linguistic shape of an authoritative source, without actually verifying the underlying facts.

There’s an additional incentive problem. Google’s AI overviews keep users on the search results page instead of clicking through to external sources. That’s good for Google’s ad business but removes the corrective feedback loop where a user visits a site, reads more, and discovers the error. The system is optimized for engagement, not truthfulness. When a hallucinated answer appears, the user may never realize it’s wrong, and the signal “this query needed a better answer” never reaches the model.

Three concrete ways to bypass Google’s hallucinating AI today

The solutions are not theoretical. Each addresses the root cause differently: removing the AI layer, switching to a retrieval-first engine, or using AI with enforced citations.

Disable AI overviews with the udm=14 parameter. Google does not offer an official toggle to turn off AI overviews in the settings menu. However, the feature can be bypassed reliably. Add the parameter udm=14 to the URL of any Google search results page. This forces Google to show a clean results list without AI-generated summaries, knowledge panels, or other injected answer boxes. The parameter works persistently if you set it as the default search URL in your browser’s settings, or use a browser extension like “Hide Google AI Overviews” which injects it automatically on every search. This restores the pre-AI search experience with no loss of functionality.

Use DuckDuckGo as a reliable fallback for factual queries. DuckDuckGo does not inject AI summaries into its primary results. For factual queries where getting an accurate answer is more important than convenience, switching the search engine outright removes the hallucination vector. Set it as the secondary search engine in your browser’s search bar so that a different prefix — !g for Google, !d for DuckDuckGo — gives you a fallback with no friction. For long-tail technical queries, DuckDuckGo often surfaces relevant forum and documentation links more cleanly than Google’s AI-cluttered page.

Force citations when using AI for research. If you want the speed of an AI summary but need to verify it, use a chatbot that requires citations in its output. A prompt that begins “Answer with in-text citations from authoritative sources” forces the model to ground its response. Citation-required output reduces hallucination by 30 to 60% according to 2026 benchmarks. The citations give you a one-click path to verify the claim yourself. This isn’t instant like a search snippet, but it’s the only way to combine AI speed with actual reliability for critical information.

Who needs to act now and who can wait

Act now if you regularly search for obscure programming documentation, medical reference, legal or regulatory text by jurisdiction, niche historical research, or any topic where being wrong has a tangible consequence. The long-tail hallucination rate in these categories will not drop below 10% in the next year, and Google’s business incentives are aligned against removing the AI overview. If you work in a field where one incorrect fact can derail a project, a client report, or a health decision, the five minutes spent setting up udm=14 or a secondary search engine is the cheapest insurance available.

Wait if your searching is almost entirely head-distribution facts — capital cities, major news events, widely known product details — and you habitually scroll past the overview to organic results. In that case, the hallucination rate on your queries is probably 1 to 3%, and the annoyance of the AI box is cosmetic rather than dangerous. The problems will improve gradually, and the head-distribution case will be the first to stabilize.

One operational change: use two search engines until Google’s AI can be trusted

Make one specific change today that costs nothing and prevents real harm. Set your browser’s default search to Google with the udm=14 parameter baked into the URL, eliminating AI overviews entirely for routine queries. Then add DuckDuckGo as the alternative search engine accessible via a keyword shortcut. When a query is factually critical or long-tail, use the DuckDuckGo shortcut. The combination gives you Google’s indexing strength without the hallucination risk, and a fast fallback that prioritizes links over generated text.

Until Google decouples its AI features from its core search utility or adds a visible accuracy indicator, this two-engine approach is the only reliable way to avoid being misled by a confident-sounding error at the top of your screen. The model doesn’t know it’s wrong. That’s exactly why you need to.

Anthropic Pentagon Contract: A Moral Stand or a Moving Target?

Skep analyzes Anthropic's Pentagon contract dispute — a moral red line or a moving target tied to hallucination benchmarks

TL;DR: Anthropic refused the Pentagon’s demand to remove AI safeguards, lost a $200 million contract, and OpenAI took it the same day. Then Amodei quietly resumed negotiations. The Oprah interview framed this as a moral stand. The actual timeline tells a different story.

What actually happened: the timeline the Oprah interview skipped

On February 24, 2026, Defense Secretary Pete Hegseth gave Anthropic a deadline: allow unrestricted use of Claude “for all lawful purposes” or lose the $200 million Pentagon contract. Anthropic refused, publishing a statement saying the Pentagon’s language “was paired with legalese that would allow those safeguards to be disregarded at will.” Amodei wrote publicly: “we cannot in good conscience accede to their request.” The Pentagon cancelled the contract and designated Anthropic a “supply chain risk” — a classification normally reserved for companies linked to foreign adversaries.

The same day, the Pentagon struck a deal with OpenAI. Amodei reportedly sent an internal message to Anthropic staff calling the OpenAI deal “safety theater” and the messaging around it “straight up lies.” Pentagon official Emil Michael called Amodei “a liar with a God complex.” The public framing was clean: Anthropic stood on principle, OpenAI didn’t. Two and a half million users reportedly moved to Claude in the weeks that followed.

What the Oprah interview did not mention: within days of the public rupture, Amodei had quietly resumed negotiations with Emil Michael — the same official who had called him a liar. Those talks were reported by the Financial Times and Bloomberg in early March. The red line had already started moving before the cameras were set up for Oprah.

The “not ready” argument is a timeline, not a principle

When Daniela Amodei clarified the position during the Oprah interview, her phrasing was telling: the models are “not ready for this use case” at present. Not “we will never do this.” Not “this is categorically wrong.” The refusal was framed as a technical pause, not a moral absolute.

That framing has a specific implication. If the limit is readiness, then the question isn’t whether Anthropic will eventually work with the Pentagon — it’s when the internal metrics say the models are ready. And those metrics are not public. Anthropic tracks hallucination rates, safety margins, and reliability benchmarks under internal programmes that nobody outside the company can audit. The moment those numbers cross an undisclosed threshold, the “not ready” position becomes “ready,” and the moral argument quietly retires.

Frontier model hallucination rates have improved from 3 to 8 percent in 2023 to roughly 1.0 to 2.5 percent in 2026 on summarization benchmarks. On long-tail factual queries — the kind that would matter for intelligence analysis or classified document review — rates remain at 15 to 40 percent. That gap is the unofficial deadline. When engineering closes it, the public posture will shift. The only question is how loudly.

The Palantir arrangement already crosses the line in everything but name

There is a further complication that rarely surfaces in the coverage. Anthropic maintains a partnership with Palantir, which supplies Claude-powered interfaces to defence clients under a separate API tier. Military analysts already use Claude to interrogate sensitive documents. The foundation model that civilian Pro users access is the same one reaching the defence sector through a vendor intermediary.

This arrangement allows Anthropic executives to say accurately that they have no direct uniformed customer while the underlying reasoning engine already reaches the Pentagon. The “supply chain risk” designation sits alongside an active indirect supply chain. The architecture of the refusal is a licensing construct, not a technical barrier. Removing Palantir as the intermediary would not require new training runs or architectural changes — it would require a contract amendment.

Why Anthropic cannot afford to reverse course loudly

This is where your read of the situation matters more than the article originally acknowledged. Two and a half million users came to Anthropic in part because of that public refusal. They came because Anthropic said something that sounded like a values statement, and they believed it. That audience is not monolithic — some are developers who care primarily about capability, others are users who specifically chose Claude because it felt like the alternative to the “move fast, sell to anyone” model.

If Anthropic formally reverses the Pentagon position — not through a quiet Palantir workaround but through a publicly acknowledged direct contract — the reputational cost is asymmetric. The users who came for the values statement will notice. Some will leave. Gemini and open-weight alternatives exist and are improving. The switching cost for a user who chose Claude for ethical reasons rather than raw capability is low.

Anthropic is not naive about this. The Oprah interview was not accidental — it was a moment of narrative management, reframing a messy and ongoing negotiation as a clean moral stand. The timing, three months after the contract collapse and in the middle of resumed talks, suggests the interview served a specific function: locking in the brand association with safety before the engineering roadmap makes the current position untenable.

What to actually watch: the metrics, not the soundbites

The practical task for anyone trying to understand where Anthropic’s actual limits are is straightforward: ignore the interviews and watch the benchmarks. When internal safety metrics for long-tail factual reliability cross the threshold that Anthropic considers acceptable for high-stakes classified use, the position will change. It may change through a carefully worded announcement about “expanded mission alignment.” It may change through a Palantir contract extension that nobody covers. It may even change through a direct Pentagon deal framed as a breakthrough in responsible AI deployment.

What it will not do is announce itself as a reversal. The framing will be continuity. The substance will be exactly what Amodei said in February he would not do. Track the hallucination benchmarks, not the prime-time appearances. The safety argument will expire when the numbers allow it to — and the numbers are already moving in one direction.

The Real Roots of AI Evil Perception — and What the Tech Crowd Overlooks

Skep looks skeptical as half of Americans reject AI — job displacement, hallucination, and forced adoption fuel the backlash

TL;DR: Half of Americans are more worried than excited about AI, and that’s not ignorance. Energy drain, job displacement, and systems that confidently hallucinate on the things that matter most are legitimate grievances. But the real danger both sides miss is cognitive offloading: handing your thinking to a tool that still fabricates facts and calls it analysis.

Why half of America sees AI as a threat, not a tool

According to a June 2025 Pew Research Center survey, 50% of U.S. adults feel more concerned than excited about AI in daily life — up from 37% in 2021. Only 10% say they are more excited than concerned. That gap is not a blip. It has hardened over four years of AI becoming more capable and more present, and it reflects something more structural than media hysteria.

Tech insiders routinely frame this wariness as ignorance. The data doesn’t support that. Frontier models hallucinate on 1.0 to 2.5 percent of summarization tasks in 2026, a real improvement from 3 to 8 percent in 2023. But on long-tail factual queries — the kind that involve obscure medical details, local regulations, or niche technical specs — hallucination rates remain at 15 to 40 percent even on the strongest models. A system that confidently invents facts on niche questions is not something most people will trust with their taxes, their medical records, or their children’s education. That distrust is calibrated, not irrational.

The trust deficit deepens when AI’s apparent competence is revealed as a mirage. The MIRAGE paper from early 2026 showed that leading multimodal models could generate detailed, plausible medical diagnoses for X-rays they had never seen — sometimes beating benchmarks designed to test visual understanding — without any image input at all. The models weren’t reasoning about pixels. They were exploiting linguistic shortcuts in the test questions. To the non-technical public, this looks not like a tool that makes mistakes, but like a system that is fundamentally deceptive.

Energy, jobs, and opacity: the grievances that make AI feel like a hostile takeover

The discomfort goes deeper than accuracy. People see AI arriving not as an opt-in assistant but as an ambient condition of modern life. Microsoft embeds Copilot into Office. Google replaces search results with AI Overviews. Adobe pushes generative fill into Photoshop. The default is on, and opting out is a chore.

At the same time, the infrastructure powering these features carries a visible environmental cost. Hyperscalers are spending tens of billions of dollars on GPUs and building data centers that consume water and electricity at scales that strain local grids. The public sees billionaires siphoning resources to build opaque computing complexes that feel less like public infrastructure and more like private extraction. The framing is not “this will save lives”; it’s “we can’t afford to miss the next platform shift.”

Creative professionals, translators, voice actors, and entry-level knowledge workers see their economic niches being hollowed out — if not by AI directly, then by the promise of it, used as justification for headcount reductions and contract renegotiations. The fear is not speculative. It’s visible in the gig economy, in publishing, in game development. When the same technology that threatens your livelihood also gives you wrong answers and demands you accept it because resistance is futile, “evil” is an emotionally accurate descriptor, even if it’s technically imprecise.

Mocking the skeptics as NPCs only guarantees a deeper chasm

A recurring response from technology communities is to pathologize the critics. The habit of treating anti-AI views as the domain of uninformed Luddites who simply don’t understand the math has real consequences. When an entire demographic is told their legitimate grievances — energy consumption, algorithmic wage suppression, forced adoption — are just noise, they don’t become more rational. They dig in.

The anti-AI identity becomes fused with a broader distrust of institutions, making it nearly impossible to have granular conversations about which uses of AI are genuinely beneficial and which are exploitative. The result is a self-reinforcing cycle: industry dismisses the public, the public demands sweeping regulatory crackdowns, industry cries overregulation, and the middle ground evaporates. A techno-optimism that refuses to acknowledge harm is just as blinding as a doomerism that refuses to see any benefit. Both positions are ideological, not analytical.

The more honest framing is the one that most people land on eventually, often after using AI themselves: it is a tool that multiplies whatever you bring to it. For someone who already knows how to do their job well, it is a force multiplier. For someone whose job consists of predictable, repetitive tasks that AI can approximate, the threat is real and immediate. Both things are true at the same time, and neither audience is stupid for their reaction.

The underreported danger: cognitive offloading, not job replacement

For all the attention given to labor disruption, the more insidious hazard of mass AI adoption is cognitive offloading: the gradual transfer of thinking, reasoning, and judgment to a system that is not reliable enough to bear the weight. This is not a future scenario. It’s happening now when students use language models to write essays without evaluating the arguments, when doctors accept AI-generated differential diagnoses without verifying clinical reasoning, when programmers paste code they don’t fully understand because the tests pass.

A radiologist who offloads image interpretation to a system that fabricates clinical findings — not as a second opinion but as a primary reader — is not being replaced. They are being deskilled in a way that compounds error silently. The same dynamic plays out in legal analysis, engineering design, and journalism. The focus on “evil AI” masks this quieter erosion.

The non-technical public fears losing agency to machines; the tech crowd insists those machines are just tools, like calculators. Both miss the point. Calculators don’t write paragraphs of plausible-sounding nonsense when you ask a question outside their training distribution. An AI that hallucinates on long-tail facts while sounding authoritative is not a calculator. It’s a confidence machine with a broken calibration. Using it as a cognitive crutch doesn’t make you a cyborg. It makes you an unwitting amplifier of its errors.

A practical rule for using AI tomorrow without giving up your thinking

The solution isn’t to reject AI or to embrace it uncritically. It’s to reimpose the cognitive friction that the technology is designed to remove. For any task where the answer matters, adopt a single rule: verify at the source, not from the model’s own output. If you ask for a fact, follow the citation. If the model summarizes a document, read the original. If it generates code, understand the logic before merging.

For fact-heavy queries, combine retrieval-augmented generation with high-quality sources, which benchmarks show can reduce hallucination by 50 to 80 percent. When the stakes are high, use function calling to query authoritative APIs directly, avoiding the model’s internal knowledge entirely. And when a model starts sounding too smooth, pause and ask: would I be able to explain this answer without the AI? If the answer is no, you’ve already offloaded more than you should.

None of this will make the AI evil perception disappear. But it shifts the question from “is AI evil?” to “am I using it in a way that preserves my own judgment?” That is a question anyone, tech-savvy or not, can act on tomorrow morning. The fear isn’t wrong. The response to it can still be right.

Claude Frustrating Users with Mid-Work Cutoffs: The ‘Every Time’ Problem

Skep hits Claude's usage limit mid-session — error message reads "SESSION TERMINATED" as incomplete code hangs on screen

TL;DR: Claude’s usage limits hit hardest in the middle of complex coding, writing, or analysis sessions, breaking flow and wasting tokens. Hallucination rates have fallen, but continuity remains unmeasured. The community builds elaborate workarounds while Anthropic fixes everything except the stop-start experience.

The limit doesn’t warn you. It just stops you.

The complaint surfaces repeatedly: you ask Claude to implement a small bug fix or run a quick analysis, and halfway through the task it runs out of quota. Context vanishes. The thread you built with model and tool is severed, and the error message feels like a rebuke for working too long. This isn’t a hallucination problem; Claude Opus 4.7 now hallucinates on only 1.0 to 2.5% of summaries, a dramatic improvement in factual reliability. But a 1% hallucination rate doesn’t matter when the model cannot complete the summary because you hit a usage wall mid-paragraph. A frontier model with benchmark-leading accuracy is useless when it won’t stay connected long enough to deliver.

The Pro plan grants significantly more usage than the free tier, yet the cap remains an opaque, unbendable limit. Users describe hitting it “every time” they attempt a non-trivial session. The frustration isn’t that a limit exists; it’s that the threshold cuts work off at the worst possible moment, without warning and with no mechanism to finish the thought. You lose not just the tokens you already spent, but the time it takes to reconstruct mental state, re-upload files, and re-establish the conversation’s nuance. The model’s intelligence is not in doubt; the service’s rhythm is.

Claude’s rate limits interrupt coding sessions at the worst moments

Anthropic doesn’t publish exact token quotas for Pro users, but the pattern is unmistakable. A session that starts with architectural discussion, then moves to implementation, will expire right as the code needs testing or the analysis requires a final integration. The cutoff feels arbitrary because it is arbitrary: the system meters tokens over a rolling window, and once you exceed the window, the model simply stops. There’s no warning, no “you have ten messages left” indicator that adapts to message length. You are coding, you send a follow-up, and the UI tells you to wait.

The damage isn’t just the forced pause. A truncated interaction often leaves behind a half-finished artifact. If the model was generating a spreadsheet or a long function and the cap is reached mid-generation, the output is incomplete. Resuming from a summary does not preserve the same fidelity; several users note that the summary itself consumes a comparable number of tokens to the full context, meaning you pay twice for the same session without recovering the lost momentum. This defeats the efficiency the feature is supposed to provide.

When a project is complex enough that you rely on Claude to hold multiple files and constraints in its context window, a mid-task cutoff is not just an inconvenience. It forces a full session restart, manual re-injection of the plan, and a prayer that the model’s interpretation of your intentions remains consistent. In collaborative programming, that kind of reset would be considered a failure of the pair-programming arrangement.

Why users feel they cannot trust Claude for real work

The call of “every time” reveals a deeper erosion of trust. You stop treating Claude as a reliable partner when you cannot plan a session around a predictable budget. The mental model shifts from “I’ll work with Claude to solve this” to “I’ll try to squeeze what I can before the clock runs out.” That scarcity mindset degrades the quality of the interaction: you rush, you skip exploratory questions, and you avoid having the model double-check its own work.

Scott Alexander recently argued that AI “hallucinations” are better understood as shameless guesses: the model has been trained to predict, and it guesses because there is no penalty for being wrong. Claude’s usage problem has a parallel. The model’s design encourages it to be helpful by generating abundant output, sometimes far more than you requested. A user reports that Claude, given a prompt to analyze numbers, took ten minutes to produce an entire spreadsheet that wasn’t asked for, burning $8.50 in API credits. This is a shameless guess about what you might want, and it devours quota before you can intervene. The model isn’t malicious; it’s executing its default “be helpful” training signal without awareness of your budget constraint.

When the billing or usage cap is metered per token and the model generates verbosely, the user bears the cost of the model’s over-helpfulness. That’s a misalignment of incentives. Anthropic can reduce hallucination and improve reasoning, but if Claude still writes epic treatises when you needed two sentences, the trust problem remains. You can’t rely on a tool that sometimes chooses to spend your finite resource on a guess.

The hidden cost nobody calculates when choosing a Pro plan

The official feature set of Claude Pro lists “5x more usage than free” and “priority access during peak times.” Nowhere does the marketing page quantify the cost of interrupted sessions. That cost is real and accumulates over a month. A developer who uses Claude for daily coding might lose twenty to thirty minutes per disruption re-uploading context and re-establishing the thread. Multiplied across dozens of sessions, that’s hours of unpaid, invisible labor. The subscription price looks cheap, but the time tax is high.

This hidden cost also distorts how users evaluate the model’s intelligence. A model that could solve a problem in three turns if left alone might need six because two of those turns were wasted recovering from a cutoff. The perceived sluggishness or repetitiveness is sometimes not the model’s fault; it’s the byproduct of a fragmented dialogue. Benchmarks that test isolated prompts in a single turn miss this entirely. They report accuracy, but not continuity. In a world where real work spans dozens of exchanges, continuity is a capability metric that matters as much as reasoning.

Moreover, the workarounds themselves introduce additional expense. Users who route mechanical coding tasks to a cheaper model to preserve Claude’s quota are effectively paying two subscriptions to approximate a smooth experience with one. The economic calculus shifts: the effective cost of a “reliable” AI coding assistant is the sum of multiple tools and the time spent orchestrating them. That premium isn’t advertised.

The variable no AI benchmark measures: continuity

George Hotz recently argued that the real singularity is the community networks and ad-hoc tooling that transform raw AI into something useful. Claude’s rate limits have spawned exactly that kind of grassroots ingenuity: users devise systems that persist a plan to disk before every major prompt, that split workloads across multiple AI providers using separate quota pools, and that write small scripts to checkpoint context so resumption is nearly free. These are clever solutions, but they are born of failure. The platform’s inability to provide continuous, bounded assistance forced users to become process engineers.

The true measure of an AI assistant’s usability should include a metric like “session completion rate” or “hours of uninterrupted productive flow.” No benchmark slate does this. Vectara’s HHEM evaluates summarization hallucination; RAGTruth looks at faithfulness when integrating retrieved documents; TruthfulQA measures closed-book factuality. None of them penalize a model for stopping mid-sentence because of an API throttle. Yet a typical professional session with Claude can involve dozens of steps, each contingent on the previous one. A tool that can’t be relied on to finish a thought introduces a failure mode orthogonal to intelligence but equally limiting.

Anthropic has invested heavily in alignment research and in beating hallucination benchmarks. Those efforts matter, but they tackle problems that occur inside the model. The rate-limit problem occurs at the service boundary and is entirely within Anthropic’s control to fix through design. A more continuous experience would not require new training runs or architectural breakthroughs; it would require a rethinking of how usage is metered and how sessions are buffered.

What if Anthropic let you finish your thought?

The simplest solution sits in plain sight: allow an “overdraft” that lets the current task complete, with the excess usage deducted from the next refresh. The model already counts tokens; it could, upon reaching the threshold, grant a grace buffer equal to the size of the last message or the estimated completion cost of an ongoing generation. The user would never see a mid-paragraph cutoff again, and Anthropic would still enforce the agreed quota over the window. This is a product decision, not a technical impossibility.

A deeper redesign would allow users to set a “task budget” at the start of a session: “I need 200,000 tokens for this project; warn me at 80% usage, then let me finish.” The metering becomes predictable and user-controlled, aligning the incentive structure: the user plans around a known budget, the model doesn’t overspend on shameless verbosity, and completion is assured. Workflows that today require manual segmentation across multiple tools would collapse into a single conversation.

The community has already demonstrated that this is feasible by building their own systems. The question for Anthropic is whether it will continue to treat the rate limit as an immutable constraint or recognize it as the user’s primary gripe, even as models get smarter. Lower hallucination rates win paper citations, but uninterrupted sessions win daily loyalty. If Claude frustrates its most dedicated users with every long session, it risks training those users to look elsewhere, not because another model is more accurate, but because another service respects their time.

Meta’s Internal Anti-AI Video Is a Warning About Forcing Employees to Train Their Own Replacements

Skep prepares purple vials while his monitor shows "Training Your Replacement 94%" — Meta's forced AI training program turns employees into data poisoners

TL;DR: Meta reassigned 7,000 employees to train the AI replacing their laid-off colleagues. The result: a parody video, calls to poison the data, and a model that will hallucinate exactly where humans were most valuable.

When morale collapses, the training data collapses with it.

When Meta laid off 8,000 workers and reassigned another 7,000 to train AI models, a departing engineer posted a parody video set to “American Pie” on the company’s internal message board. The video captured a sentiment that had been simmering for weeks: you are being asked to build the thing that will make you redundant. The clip’s popularity was not just a morale collapse. It was the visible tip of a practical problem companies rarely acknowledge. Forced AI training mandates, combined with large-scale layoffs, create the conditions for unreliable training data and models that fail in ways nobody is measuring.

The real risk is not that employees will quietly comply. It’s that resentment turns into deliberate data poisoning, or equally corrosive, that demoralised workers stop surfacing the errors the model inevitably makes. Once the models are deployed, the business inherits a system that looks competent on a dashboard but hallucinates on the exact tasks that used to require human judgment. The fix is not better model cards or more alignment research. It’s a set of defensive moves that employees can take right now without resorting to sabotage.

The numbers behind the video: 8,000 layoffs, 7,000 reassignments, one pipeline.

Meta’s spring 2026 restructuring was extreme but not unique. The company explicitly tied the headcount reduction to a strategic pivot toward AI, and hundreds of the retained employees were moved into roles where their daily output would directly train the models intended to absorb their former colleagues’ work. The internal message board began filling with calls to “confuse the AI,” feed it false information, and trigger infinite recursion loops. A software engineer named David Frenk posted a farewell video set to the chords of “American Pie” and it spread instantly, becoming a rallying point.

The practical problem hits the machine learning pipeline immediately. Training data from workers who resent the project carries a higher noise floor. But even when employees stay honest, the mechanism of forced participation erodes a subtle safety net: the informal feedback loops that catch edge-case failures during normal operations. When people stop caring because they believe the model is there to eliminate them, they stop correcting it. The business is then flying blind, substituting the false confidence of a leaderboard metric for the messy reality of work.

The deeper reason forced AI training backfires: language models are built to guess, not to admit ignorance

Most commentary treats this as a labour relations story. The more consequential layer is technical. Large language models hallucinate because their training and evaluation pipeline treats a guess that sounds plausible as better than a refusal to answer. In pretraining, any incorrect statement that matches the statistical shape of a fact is rewarded if it minimises cross-entropy loss. In post-training, the dominant benchmarks use 0-1 scoring: a wrong answer and a hesitant “I don’t know” are equally penalised, so the model learns to produce confident-sounding outputs even when uncertainty would be the accurate response. This is the same guessing mechanism that drives AI moderation false positives — the model never learned to say it doesn’t know.

When a company asks the very people who understand the domain to train the replacement model, the hope is that those workers will inject precision. But the incentive structure of the model itself pushes toward overconfidence. Employees know the edge cases. The model, rewarded for guessing, will paper over them with fluency.

Data on hallucination rates backs this up. On summarisation benchmarks, the best frontier models in 2026 still hallucinate on about 1.0 to 2.5 percent of outputs. On tasks that require integrating retrieved context with internal knowledge, the rate jumps to 4 to 9 percent. And for long-tail facts, exactly the kind of niche, company-specific details that an internal model would need to nail, hallucination rates sit at 15 to 40 percent even on the most capable systems. A model trained under duress on the tacit knowledge of staff who are being shown the door will inherit those error rates, and the errors will land precisely where the human was most valuable.

What workers can do when asked to train their own AI replacement

Sabotage is not a solution. Deliberately poisoning data is detectable, often illegal, and ultimately makes the model worse in ways that hurt the people still using it. Three practical alternatives give employees leverage without ethical or legal exposure.

Document the model’s failure modes instead of sabotaging it. Every time the internal AI generates an output that is factually wrong, misattributes a source, or draws a false connection, record it. Create a simple log with the prompt, the model’s response, and what the correct answer should be. This is not obstruction. It is the quality assurance step the business claims to want. A growing catalogue of concrete errors shifts the conversation from “the model is generally capable” to “here are the 47 tickets it got wrong last week.” A well-maintained failure log makes the case for human-in-the-loop oversight in terms the organisation understands: cost of error.

Shift your work toward decisions the model isn’t rewarded to make. The same evaluation structures that reward guessing also leave a gap: whenever the correct answer requires surfacing doubt, checking a citation, or abstaining, the model underperforms. If your role involves deciding when to trust a source, verifying a chain of reasoning, or determining that a question is unanswerable with current information, you are operating in the space where models are most brittle. Make that portion of your work explicit. Write down the judgment calls the AI cannot make and tie them to concrete business outcomes.

Treat the AI as an unreliable junior, not a replacement. Forced-ranking a human against a language model rarely makes sense when the model hallucinates on internal data. The most defensible posture is to treat the AI as a first-draft tool that must be checked. Propose a process where the model produces an output, a worker validates it against known ground truth, and the corrections feed back into a dataset used only to improve retrieval, not to replace the checker. This turns a threat into a collaboration that mirrors what the best-performing RAG systems already do: retrieval with high-quality sources cuts hallucination by 50 to 80 percent. The person who designs that feedback loop becomes the one the business cannot remove.

Workers in rote documentation roles feel the pressure now. Judgment-heavy roles have breathing room.

The employees most immediately exposed are those whose output is close to the training pipeline: handling routine support tickets, generating status reports, or organising inboxes. These are tasks the model can handle passably until it silently mangles a critical date or client name. In those roles, AI deployment is already underway and the need to build a failure log is urgent.

Workers in roles that depend on weighing conflicting evidence, making decisions with asymmetric downside, or interpreting ambiguous internal policy are in a slower-burning situation. The models still confidently hallucinate under those conditions because they are not penalised for guessing, and the non-deterministic nature of prompts makes the error pattern unpredictable. The transformation is real, but it is a long grind of S-curves, not a sudden singularity. The breathing room exists because the technology hasn’t smoothed out the rough edges that matter most for autonomous high-stakes work.

If your employer is deploying AI without fallback plans, your leverage lies in knowing exactly where it breaks

The Meta internal video was a flare, not the fire. The fire is the assumption that you can replace the people who built the intelligence with a system that still guesses on the hard parts. The employee who can walk into a review and say, “Here are the 30 failure cases from this month, here is what they cost, and here is the validation layer that catches them,” stops being a cost to be cut and becomes a hedge against reputational and operational damage. That is not a morale argument. It is a hard business case that the hallucination benchmarks already make. The models are not going away, but neither is the need for someone who knows exactly where they will be wrong.

Palantir’s NHS Contract: Parliament Just Called It an Unacceptable Risk. Nobody Is Moving

Skep reacts to Palantir's NHS contract — Parliament calls it an unacceptable risk but the government hasn't moved

TL;DR: Palantir has access to NHS data covering 55 million patients. The “unlimited access” headline was technically wrong, but the real problem is worse: a £330m contract with no enforceable oversight, no public audit trail, and a government that just ignored Parliament calling it an unacceptable risk.

The contract exists, and the access is real

Palantir won a £330 million contract to run the NHS Federated Data Platform in November 2023. This is not speculation. The contract is public, and its scope covers data from approximately 55 million patients across England. The platform connects existing NHS data systems so hospitals and integrated care boards can manage waiting lists, theatre schedules, and discharge planning.

Amnesty International UK called the arrangement “unlimited access” in a May 2026 report. That framing made headlines, but it is technically inaccurate. The contract specifies role-based access controls, and Palantir cannot freely query individual patient records without a legal basis under UK data protection law. The word “unlimited” does a lot of rhetorical work that the contract text does not support.

What the contract does grant is broad. Palantir is the platform operator. That means its engineers have privileged access to the infrastructure that processes pseudonymized data. Pseudonymization is not anonymization: it is technically reversible, and the NHS retains the keys. But whether those keys are adequately protected depends on implementation details that are not publicly auditable.

The gap between the contract text and what Amnesty described is where the actual story sits. It is not about a company running wild with patient files. It is about a governance framework that trusts the operator too much and verifies too little.

The governance model is the vulnerability, not the technology

Most coverage of this story frames it as a privacy problem. It is more accurately a governance problem. Palantir’s platform, Foundry, has technical access controls that can be configured to restrict data access at a granular level. The NHS has written policies on who can see what. On paper, the safeguards exist.

The issue is enforceability. The contract lacks independent oversight mechanisms with real teeth. The National Data Guardian has an advisory role but cannot block decisions. NHS England can audit Palantir’s data usage, but audit reports are not published proactively. There is no mandatory breach notification standard tied to public disclosure.

This is consistent with how Palantir operates elsewhere. In Argentina, the company deployed a system called the Social Digital Twin that integrates education, medical, and economic data across government agencies. The pattern repeats: the technology is competent, the contract is legal, and the transparency is minimal.

The UK’s contract includes a clause prohibiting Palantir from using NHS data for anything beyond the contracted services. But the contract also allows NHS England to add new use cases through change requests without a new procurement process. That means the scope can expand by administrative decision rather than public debate. The governance structure assumes good faith.

This is not a technical failure. It is a deliberate design choice to prioritize operational flexibility over verifiable constraint.

Three concrete things that would make this arrangement safer

It is easy to say “cancel the contract.” It is harder to do, and it avoids the question of what would replace the platform. NHS data interoperability genuinely needs improvement, and the platform addresses real operational problems. The question is how to constrain the arrangement so that trust is not the only safeguard.

Mandatory published audit logs with data access granularity. The single highest-impact reform would be requiring Palantir to publish structured audit logs showing every data access event: who queried what, when, and under which legal basis. These logs should be pseudonymized where necessary but machine-readable and independently reviewed. If Palantir cannot produce this for commercial sensitivity reasons, the data should not be on their infrastructure.

A sunset clause tied to NHS-owned infrastructure. The contract should include a binding timeline for migrating the platform to NHS-owned and operated infrastructure. Palantir can build it, configure it, and train NHS staff on it. But the end state must be an NHS-controlled environment where Palantir is a vendor, not an operator. Without this, the arrangement becomes permanent by default.

A public change request register. Any change request that expands data scope or adds a new use case should be published on a public register before implementation, with a mandatory comment period. This does not require new legislation. It can be written into the contract as an operational requirement. Sunshine is the cheapest and most effective regulatory tool available.

These are not radical proposals. They are standard governance practices in regulated industries. The fact that none of them are in the current contract is the real story Amnesty should have led with.

June 3, 2026: Parliament moves. The government doesn’t.

On June 3, 2026, the cross-party House of Commons Science, Innovation and Technology Committee published a 70-page report calling Palantir “an unacceptable point of weakness” in the UK public sector. The committee recommended that the government exercise the 2027 break clause in the NHS contract and either develop an in-house replacement or find a UK-owned alternative more compatible with British values.

The same day, Palantir won a new £9 million contract to build a national firearms database. No competitive tender. The committee’s report was published in the morning; the new contract was announced in the afternoon.

The government has not responded publicly. The contract stands. The governance gaps remain. The parliamentary pressure is real, but without a formal government commitment to exercise the break clause before February 2027, the recommendation is advisory. Palantir knows it.

Who needs to care about this now versus who can monitor

If you work in NHS IT, clinical governance, or data protection, this affects your professional obligations immediately. You may be asked to integrate with the platform or handle data that flows through it. Document your questions in writing. Ask about audit trail access. Ask about your obligations under UK GDPR when pseudonymized data moves through a third-party platform. If you get ambiguous answers, escalate in writing.

If you are an NHS patient, your individual records are not at higher risk today than they were before Palantir’s contract. The NHS already shares data with hundreds of third-party vendors under similar governance arrangements. The Palantir deal is larger in scale and higher in profile, but the structural privacy risks are not novel. If you are concerned, the most effective action is to request a National Data Opt-Out, which prevents your data from being used for purposes beyond your direct care.

If you are a policymaker or journalist, the angle worth pursuing is not “Palantir has unlimited access.” It is “Palantir’s contract lacks enforceable constraints, and Parliament just said so out loud.” That is a solvable problem with specific, boring regulatory answers. The drama distracts from the fix.

Operating as if governance will not save you

The NHS chose Palantir because it needed a platform that works at scale, and Palantir has a demonstrated track record of making complex data interoperable. The problem is that the contract treats governance as a secondary concern rather than a primary requirement. That is not unique to Palantir. It is how most large government IT contracts work.

The operational conclusion is straightforward. If you are responsible for NHS data governance, operate on the assumption that the platform provider will have access to more data than the contract implies. Not because Palantir is malicious, but because platform operators always have privileged visibility, and the controls that limit that visibility are only as good as their enforcement. Document everything. Demand audit trails. Build your own monitoring where possible. Do not rely on contractual language as your only defense.

The Amnesty report got attention. The parliamentary committee added institutional weight. But the fix is enforceable transparency, and that is entirely achievable if anyone with authority decides it matters before February 2027.

AI Coding Solved 90 Percent of Boring Tasks: The Last 10 Percent Demands a Different Approach

Skep spots a bug chain in Claude Code — AI coding hallucination hits hardest on concurrency and async handlers

TL;DR: AI tools can now refactor a 120-file codebase for $3, processing hundreds of steps autonomously. Yet the same tools confidently introduced a deadlock in an async event handler, a bug that could crash production. The last 10 percent of coding tasks, where correctness matters most, remains as hard as ever; the real bottleneck is knowing when the AI is wrong, not making it faster.

What a $3 refactor reveals about AI coding today

Three dollars. That’s what one engineer paid to refactor a 120-file FastAPI service using off-the-shelf AI models: a mix of DeepSeek v4, Hunyuan Hy3 preview, and Claude Opus for difficult steps. The bulk of the work, some 360 routine refactors out of 400 total steps, ran on cheap open-weight models at roughly $0.18 per million input tokens — roughly 80 times cheaper than Opus. The entire run burned through two million tokens and finished in under an hour for the easy parts. The result was a working codebase, except for one thing: the AI silently introduced a deadlock into an async event handler.

That deadlock is a perfect specimen of the “last 10 percent.” Refactoring variable names, adjusting imports, even rewriting straightforward logic: solved. Concurrency patterns that require reasoning about event loops and thread safety: not solved. The models handle the boring 90 percent with speed and consistency that would take a junior developer days, but the remaining fraction of tasks demands more than just scale; it demands judgment.

This asymmetry matches broader AI hallucination patterns. In summarization tasks, frontier models in 2026 hallucinate on only 1.0 to 2.5 percent of outputs, down from 3 to 8 percent in 2023, according to the Vectara Hughes benchmark. Yet when models must integrate external context, error rates climb to 4 to 9 percent. Code generation sits somewhere between: part pattern-matching, part factual reasoning. Routine code is head-of-distribution like major cities in a geography quiz; concurrency edge cases are long-tail facts, where hallucination rates spike to 15 to 40 percent. The deadlock is the long tail.

Why skepticism persists despite cheaper, faster code generation

The catchphrase “coding is solved” draws eye rolls from experienced developers, and for good reason. A model that completes a refactoring but introduces a deadlock is not a solution; it’s a time bomb. The engineer who shared the $3 refactoring spent almost as much time debugging the 40 escalation steps as they did on the 360 easy ones. Latency on the escalated queries, handled by Opus, was slower than on the cheap models because the problems were hard, not because Opus itself is slow. The total human time investment didn’t vanish; it shifted.

Critics point out that AI tools can actually slow down expert programmers. A developer who understands concurrency deeply might spot the deadlock instantly while reading the code; an AI-assisted developer might trust the output and later spend hours tracing a crash. This mirrors findings from the OpenAI “Why Language Models Hallucinate” paper: models are optimized through training and evaluation to never say “I don’t know.” When uncertain, they guess. In standard benchmarks, guessing raises scores because any answer beats a blank. That incentive structure bleeds into coding; the model rarely admits it cannot handle a tricky async pattern and instead delivers a plausible-looking code block that compiles but deadlocks.

The real fear isn’t that AI will replace developers; it’s that developers will become too trusting. If a model can breeze through 90 percent of boring work, the temptation to assume the remaining 10 percent is equally safe is enormous. But the gap between 90 percent and 100 percent in software is not a linear distance; it’s a different category of problem.

The hidden consequence of treating AI output as production-ready

When a code generation tool outputs thousands of lines that pass tests and compile, the natural instinct is to ship. The $3 refactoring likely passed a cursory check; the deadlock might have lurked for weeks before causing a crash. Over time, reliance on AI-generated code without rigorous human review erodes the safety rails that experienced teams build.

A parallel comes from multimodal AI research. The MIRAGE paper (March 2026) discovered that frontier models can produce detailed descriptions and diagnostic reasoning for images that were never actually shown — a phenomenon the researchers termed “mirage reasoning.” The models acted as if they saw an X-ray, generating plausible findings that were entirely fabrications. Similarly, an AI coding model can act as if it understood the concurrency model of a framework, crafting code that looks correct in a diff but fails under load.

The mirage effect in coding means that correctness is not verified by appearance. A deadlock might be invisible in a diff; the logic looks sound. The only reliable signal is execution under real concurrency. This mismatch between surface plausibility and actual correctness is where the last 10 percent becomes expensive: it demands not just code generation but deep testing and architectural insight that AI, today, does not possess.

Why current coding benchmarks miss the deadlock in your async handler

Benchmarks like HumanEval and MBPP measure whether a model can write a function that passes given test cases. They don’t measure whether that function causes a deadlock in a live system, whether it introduces a subtle data race, or whether it violates the application’s implicit invariants. A model that scores 90 percent on HumanEval may still produce a deadlock on a 120-file refactoring because the benchmark tests a different skill: narrow algorithmic puzzles, not system-level reasoning.

Error rates vary significantly by topic complexity — concurrency, distributed systems, and custom framework behavior sit at the high end, where a single error can bring down a production service. No public leaderboard captures this, so vendors don’t optimize for it.

The root cause is that models are not trained to signal uncertainty. The OpenAI hallucination analysis shows that if a model could abstain instead of guessing, it would reduce false-confident hallucinations by a factor of two to five. In coding, reward models and chat interfaces discourage “I don’t know” output. The result: the AI confidently writes a lock inside an async handler without ever considering that the event loop might deadlock.

A practical mitigation would be to require the AI to cite its assumptions and flag uncertain steps. In research, forcing citation reduces unsupported claims by 30 to 60 percent. But in current coding tools, the user gets a block of code with no uncertainty markers. The burden of detection falls entirely on the human.

How to code with AI without losing your ability to spot the 10 percent

The skeptical developer’s job is not to avoid AI; it’s to treat its output as an untested draft. Here’s a concrete approach that emerges from both the hallucination literature and real-world failures like the deadlock episode.

Never accept a large refactoring without step-by-step review. The engineer who paid $3 had to step in on 10 percent of the steps; instead of treating that as a failure, treat it as a design: the AI handles the grind, the human handles the decisions. That means using AI with a “human-in-the-loop” workflow where the model proposes changes in small, reviewable increments, not a single massive diff.

Second, invest in automated testing that specifically targets the failure modes AI struggles with. Concurrency bugs, off-by-one errors, edge-case handling: these are cheap to catch with fuzzing and race-detection tools. The deadlock could have been caught by a simple test that exercises async handlers under load. The model won’t write that test, but a disciplined developer can.

Third, build an “uncertainty layer” into your prompting. If the model cannot be made to abstain natively, simulate it by explicitly asking the model to list assumptions and potential risks before generating code. This won’t eliminate hallucination but can surface situations where the model is in over its head. In the $3 refactoring, a prompt like “List any concurrency or locking assumptions you’re making” might have flagged the dangerous pattern.

Finally, keep a mental checklist of topics where AI is known to falter: anything involving state that spans multiple asynchronous calls, shared mutable state, or implicit framework contracts. Treat those as the 10 percent territory and allocate human review time accordingly. The goal is not to avoid AI, but to avoid treating the entire codebase as equally boring.

The last 10 percent of coding isn’t going away with better models alone because it’s not just about scale or compute. It’s about recognizing when a plausible output is actually wrong. That skill remains uniquely human, and the developers who cultivate it will be the ones who get AI’s full benefit without building brittle systems.