Is ChatGPT Really Getting Worse? (2026 Facts)
I’ve spent over three decades in IT, and one pattern I’ve learned to trust is this: when a large number of experienced, level-headed users independently report the same problem with a tool they use daily, it’s worth investigating properly rather than dismissing it as noise. Is ChatGPT Really Getting Worse? Separating Perception From Fact is exactly that kind of investigation, and I went looking for actual controlled research rather than more forum speculation.
Is ChatGPT Really Getting Worse? Separating Perception From Fact means using controlled, repeatable testing to check whether a model’s actual measured performance changed over time, rather than relying on memory or gut feeling. For example, a Stanford and UC Berkeley study found GPT-4’s accuracy on a prime number identification task dropped from 84.0% to 51.1% between two versions released just months apart.
If you’ve felt like ChatGPT quietly got worse and worried you might just be imagining it, I want to address that directly before anything else: you’re not alone, and in a meaningful number of documented cases, you’re not wrong either.
Is ChatGPT Really Getting Worse, or Does It Just Feel That Way?
Quick Answer
Yes, ChatGPT’s behavior demonstrably changes over time — this isn’t just perception. A peer-reviewed Stanford and UC Berkeley study found GPT-4’s accuracy on identical math tasks dropped by over 30 percentage points between its March and June 2023 versions, while GPT-3.5 actually improved on the same tasks. OpenAI itself acknowledged some tasks can get worse even as most metrics improve overall.
In my experience, the hardest part of this problem for most users isn’t the decline itself, it’s the uncertainty around whether they’re allowed to trust their own observation. I’ve seen plenty of forum threads where someone raises a legitimate concern about model drift and gets told they’re imagining it. The research says otherwise.
What Does the Research Actually Show About Model Drift?
Here’s where I want to slow down and walk through the actual numbers, because vague claims of “it feels worse” don’t hold up to scrutiny the way specific measured figures do.
GPT-4’s Math Accuracy Dropped Over 30 Percentage Points
On a prime number identification task, GPT-4’s accuracy fell from 84.0% in March 2023 to 51.1% in June 2023, using the exact same set of questions run against both model versions. Harvard Data Science Review
That’s not a marginal wobble you’d expect from ordinary run-to-run variation — that’s a collapse of over 30 percentage points on a task the earlier version handled comfortably.
Response Length Collapsed on the Same Task
Response verbosity on that task dropped from an average of 638.3 characters to just 3.9 characters, meaning the newer version frequently answered with a bare “No” instead of showing any reasoning. Harvard Data Science Review I find this detail more revealing than the accuracy number alone, because it points at a specific behavioral shift rather than just random noise: the model stopped explaining itself, even when explicitly asked to.
GPT-3.5 Improved While GPT-4 Declined on the Same Tasks
GPT-3.5’s accuracy on the identical prime number task rose from 49.6% to 76.2% in the same window. Scientific American I want to flag this specifically because it kills the simplest explanation — this isn’t a story about AI models getting worse in general. It’s a story about specific capabilities shifting in specific directions on specific models, sometimes for the better and sometimes for the worse.
Why Does ChatGPT’s Behavior Change Between Versions?
Understanding the mechanism matters more than just accepting the numbers, because it changes what you should actually do about it. Here’s the sequence I walk through with anyone trying to figure out whether their workflow is at risk.
- Confirm the drift is real, not imagined. Peer-reviewed research directly measured accuracy swings of over 30 percentage points on identical tasks run months apart on the same model name.
- Identify which pattern matches your complaint. Shorter responses point to a verbosity drop, more refusals point to safety tuning, and worse math or logic performance points to declining chain-of-thought following.
- Re-verify chain-of-thought prompting still works. Researchers found that explicitly asking a model to “think step-by-step” can silently stop producing intermediate reasoning between model versions, even with an unchanged prompt.
- Pin to a dated model snapshot for critical workflows. Use a specific dated API model version rather than a rolling default name whenever your workflow depends on consistent behavior.
- Re-test prompts periodically, not just once. Researchers explicitly recommend continuous monitoring of LLM behavior as an ongoing practice rather than a one-time setup task.
- Understand that gains and losses happen together. A model can become measurably safer and more resistant to jailbreaks while simultaneously getting worse at math, so “is it getting worse” has no single answer without specifying the task.
- Take forum complaints seriously as a signal. Researchers built this exact study in direct response to widespread but previously unverified user complaints about ChatGPT feeling worse.
I want to expand on step 3, because it’s the one I see catch experienced developers off guard most often. Chain-of-thought prompting — the practice of asking a model to work through a problem step by step before giving a final answer — is one of the most reliable ways to improve reasoning accuracy. The researchers documented a case where the March 2023 version of GPT-4, asked whether 17,077 is prime and told to think step-by-step, produced a full four-step reasoning chain and reached the correct answer. The June 2023 version, given the identical prompt, skipped the reasoning entirely and answered “No” — which was incorrect. Harvard Data Science Review
Prompt: Is 17,077 a prime number? Think step by step.
March 2023 GPT-4: [Full 4-step reasoning chain shown] -> Correct answer
June 2023 GPT-4: "No" -> Incorrect answer, zero reasoning steps shown
That’s a real documented finding from the study, not an illustrative example. If your workflow depends on step-by-step reasoning actually appearing in the output, this is precisely the kind of silent behavioral shift that can break things without any error message at all.
What Does OpenAI Itself Say About This?
I always want to check whether the company involved has said anything on the record, because a third-party study is convincing but an admission from the source is stronger still. OpenAI acknowledged this pattern directly in a July 20, 2023 blog update, stating: “while the majority of metrics have improved, there may be some tasks where the performance gets worse.” Scientific American
That’s a fairly direct confirmation of the core finding. OpenAI isn’t claiming universal decline, and the research doesn’t either — but both sources agree that improvement isn’t uniform across every task, every time a model gets updated.
What’s Actually Driving This Behavior Shift?
The researchers’ leading explanation is that GPT-4 lost some of its ability to reliably follow explicit chain-of-thought instructions between the two versions tested, and that this single change explains much of the accuracy collapse observed on reasoning-heavy tasks. Harvard Data Science Review In my own reading of how these systems get tuned over time, this tracks with a broader pattern I’ve seen across the industry: models get adjusted for RLHF safety tuning, response conciseness, and inference cost optimization, and those adjustments can have side effects on unrelated capabilities that nobody explicitly intended to change.
I’d stop short of assuming malice or a deliberate quality cut here. The far more mundane explanation — and the one the evidence actually supports — is that optimizing a model for one goal, like safety or shorter, cheaper responses, can unintentionally degrade performance on a completely different axis, like careful multi-step arithmetic. That’s a real engineering tradeoff, not a conspiracy.
Is Every Part of ChatGPT Actually Getting Worse?
No, and this is the part that gets lost in most casual complaints about the tool. The study is explicit that GPT-4 became measurably safer over the same period, answering fewer sensitive or dangerous questions and resisting jailbreak attempts more successfully. Scientific American So if your specific complaint is about math accuracy or reasoning depth, the research backs you up. If your complaint is that the model has become more cautious or refuses more requests, that’s also documented — but it’s a different phenomenon with a different underlying cause.
I think this distinction matters more than people initially realize, because lumping every complaint under one umbrella of “ChatGPT is getting worse” makes it impossible to have a precise conversation about what actually changed. Benchmark regression on one task and increased caution on another are not the same problem, and they don’t share the same fix.
Bad vs. Good Way to Think About ChatGPT Feeling Worse
Let’s put the two framings side by side, because how you interpret this experience determines whether you waste time or actually solve your problem.
Bad: “ChatGPT keeps giving me one-line non-answers now, this must just be my imagination or bad luck since OpenAI would never let quality drop.”
Good: “I noticed ChatGPT’s math and coding answers got noticeably shorter and less reasoned over a few months, and I confirmed this matches a peer-reviewed study that measured GPT-4’s accuracy on identical tasks dropping over 30 percentage points between two model versions, so I’ve started pinning my API calls to a specific dated model snapshot instead of the rolling default.”
The bad framing treats a documented, measurable phenomenon as a personal delusion, which leads nowhere productive. The good framing treats it as a known risk with a known mitigation, which is exactly what it is.
How Should You Actually Test This Yourself?
If you want to verify this for your own use case rather than taking any study’s word for it, the approach the researchers used is straightforward enough to replicate on a smaller scale. Save a fixed set of representative prompts from your actual workflow, along with the outputs you got at the time.
Periodically re-run that exact same set of prompts and compare the new outputs against your saved baseline. Look specifically for changes in accuracy on tasks with a clear right answer, changes in response length or level of detail, and changes in whether the model shows its reasoning when asked to. The official GitHub repository behind the Stanford and UC Berkeley study actually publishes the raw prompts and data they used, which is a genuinely useful reference if you want to see exactly how a rigorous version of this test looks in practice. GitHub
What Should You Do If Your Workflow Depends on Consistent Output?
If you’re running ChatGPT or the API in any kind of production workflow, automated pipeline, or business process, this research has a direct practical implication you shouldn’t ignore. Relying on a rolling default model name means your workflow’s behavior can shift underneath you without any announcement, any error, or any warning.
I’d recommend pinning to a specific dated model snapshot wherever the API supports it, precisely because that gives you control over when an update actually reaches your system. Treat model updates the way you’d treat a dependency upgrade in any other software project — something you opt into deliberately after testing, not something that happens to you automatically in the background.
How Often Should You Re-Test Your Prompts?
I don’t think there’s a single universal answer here, but I’d suggest treating it the way I’ve treated regression testing throughout my career: tie it to events, not just a calendar. Re-test whenever you notice a provider announcing a new model version, whenever your own results start feeling inconsistent, and at minimum on a quarterly basis for anything business-critical.
The researchers’ own recommendation leans in this exact direction — they explicitly call for continuous monitoring of LLM behavior rather than a one-time evaluation, precisely because a model that passes your test today offers no guarantee about how it will behave after its next update. Harvard Data Science Review
| Signal You’re Noticing | Likely Cause | What The Research Found |
|---|---|---|
| Shorter, less detailed answers | Response verbosity drop | Average length fell from 638.3 to 3.9 characters on one measured task |
| More refusals or caution | RLHF safety tuning | GPT-4 became measurably better at resisting jailbreaks and sensitive requests |
| Worse math or logic accuracy | Declining chain-of-thought following | Accuracy dropped over 30 percentage points on identical prime number questions |
| Inconsistent results across sessions | Normal run-to-run variation, separate from drift | Not the focus of this study, but a distinct and equally real phenomenon |
Does This Pattern Show Up With Other AI Providers Too?
It’s worth asking whether this is a ChatGPT-specific quirk or a broader pattern across the AI industry, since that changes how you should think about the risk going forward. The specific study I’ve referenced here focused on GPT-3.5 and GPT-4, but the underlying mechanism — providers updating a model behind a stable-sounding name, sometimes with unintended side effects on specific capabilities — isn’t unique to any single company.
Any provider that ships continuous updates to a model without versioning every release as a distinct, permanently addressable snapshot creates the same basic risk. I’d treat this as a structural feature of how commercial AI services are currently built and maintained, not a flaw specific to one company, which is exactly why the mitigation — pinning to dated versions and re-testing periodically — applies regardless of which provider you’re using.
What Should Casual Users Take Away From This?
Not everyone reading this is running a production pipeline or an automated testing suite, and I don’t want to overstate the urgency for casual, everyday use. If you’re using ChatGPT for quick questions, drafting emails, or general conversation, the practical impact of this kind of drift is usually minor and rarely worth restructuring your entire workflow around.
Where I’d actually pay attention as a casual user is if you notice a sudden, specific change in a task you rely on regularly — say, a coding assistant that used to explain its logic now just dumps code with no explanation, or a research assistant that used to double-check its math now gives confident but wrong answers. Those are the moments worth pausing on, checking whether the pattern matches what the research describes, and adjusting your expectations or your prompting approach accordingly, rather than assuming either that nothing has changed or that everything is now broken.
For a broader look at troubleshooting common AI tool problems and understanding how these systems actually behave, see our complete guide to troubleshooting AI tools.
Frequently Asked Questions
Is ChatGPT really getting worse, or is it just perception?
It’s real and measurable — a peer-reviewed study found GPT-4’s accuracy on identical math tasks dropped over 30 percentage points between two versions released months apart. Harvard Data Science Review
Why did GPT-4 start giving shorter answers?
Researchers measured average response length on one task collapsing from 638.3 characters to 3.9 characters between versions, suggesting a shift toward terser outputs on that specific task.
Does this mean every part of ChatGPT is getting worse?
No — GPT-3.5 improved on the same tasks in the same period, and GPT-4 became measurably safer against jailbreak attempts, showing decline isn’t uniform across all capabilities. Scientific American
Has OpenAI acknowledged this issue?
Yes — OpenAI stated in a July 2023 update that while most metrics improve with new versions, some tasks can get worse. Scientific American
How can I protect my workflow from unexpected model changes?
Pin your API calls to a specific dated model snapshot rather than a rolling default name, and periodically re-test critical prompts rather than assuming past performance holds indefinitely.
Why did ChatGPT stop showing its reasoning steps?
Researchers found that chain-of-thought prompting can silently stop working between model versions, causing the model to skip intermediate reasoning even when explicitly asked to think step-by-step. Harvard Data Science Review
Is there a way to actually test this myself instead of just trusting a study?
Yes — save a fixed set of prompts and outputs from your own workflow as a baseline, then periodically re-run the same prompts and compare accuracy, length, and reasoning depth against that baseline.
Leave a Reply