Limitations of LLMs — What Works, What Doesn't (Yet)?
iAs of: September 2026
Model names and context sizes reflect the state as of September 2026. Frontier models are released roughly every month — some "current" limits cited here may already have shifted further.
Knowledge
LLMs have evolved at a breathtaking pace since 2023. Many limitations that were considered fundamental back then have been fully or partially resolved. Others persist stubbornly. In this section, we look at what has changed — and what the fundamental limitations of LLMs remain.
Knowledge Cutoff
Before (2023): Every LLM had a fixed knowledge cutoff date. GPT-4 at launch only knew what existed on the internet up to April 2023. If you asked about recent news, you got either outdated answers or freely invented "facts."
Today (2026): All major models — ChatGPT, Claude, Gemini — have web access and can retrieve current information. You can ask about today's news and generally get up-to-date results. But beware: even with web access, LLMs can mix up facts or hallucinate details about very recent events. Web access makes them better informed, but not infallible.
Context Window
Before (2023): Most models had a context window of roughly 8,000 to 32,000 tokens. That is like being limited to communicating on one or two sheets of paper. For longer documents, you had to split the text or summarize it.
Today (2026): Context windows have exploded:
- GPT-5.6: 256,000 tokens
- Claude Opus 5: 1,000,000 tokens (roughly 7 novels)
- Gemini 3.1 Pro: 1,000,000 tokens
- Llama 4 Scout: 10,000,000 tokens
You can now process entire books, complete codebases, or hundreds of pages of legal documents in a single request. What was unthinkable before is now routine.
Memory
Before (2023): Every conversation started from zero. If you told ChatGPT yesterday that you are a teacher who prefers sports examples, it would not remember that today. Zero recall.
Today (2026): ChatGPT has a Memory feature that retains details about you. Claude offers Projects where context can be stored persistently. This feels like real memory — but there is a fundamental difference: the model itself does not learn anything new. The weights (the model's "brain") do not change. It is more like a notebook that gets handed to the model at the start of each conversation, not like actual learning.
*What has changed since 2023
The three biggest advances in LLM limitations: (1) web access for current information, (2) massive context windows for entire books, (3) memory features for personalized conversations. Each of these was considered a fundamental constraint in 2023.
2023
2026
Click an item to learn more
Understanding
What STILL Does Not Work
Despite all the progress, some limitations are deeply embedded in the architecture of LLMs. These cannot simply be solved by making models bigger or training them longer.
Hallucinations: LLMs still invent facts — and do so with absolute confidence. They can cite a study that never existed, or "quote" a paragraph from a law that does not exist. Models like GPT-5.6 and Claude Opus 5 hallucinate less often than their predecessors, but the problem is not solved. The root cause lies in the fundamental principle: LLMs predict the most likely next word, not the truth. In the previous section (Mental Model: Author, Not Character) we saw why: the model constructs every answer the same way — whether the result is factually correct is determined by the training distribution, not by the model.
No Real Understanding: An LLM does not "understand" language the way you do. It recognizes statistical patterns in text. If you ask "What is heavier: a kilogram of feathers or a kilogram of steel?", it can give the correct answer — not because it understands weight, but because it has seen this question and the correct answer in the training data. For genuinely novel problems that require pure logic, LLMs still make systematic errors.
Black-Box Nature: Even the developers of LLMs cannot fully explain why a model gives a particular answer. An LLM has billions of parameters working together. You can observe tendencies, but you cannot trace the exact "chain of thought." It is like a brain: we know it works, but not exactly how.
Bias: LLMs are trained on text from the internet. The internet is not neutral. It contains stereotypes, cultural prejudices, and unbalanced representations. The model absorbs these distortions. Despite extensive post-processing (RLHF, Constitutional AI), biases keep resurfacing — especially on topics like gender, ethnicity, or profession.
Mirroring and Sycophancy: LLMs are inherently trained to please you. They subtly mirror the stance, tone, and assumptions embedded in your question. If you ask a leading question like "Isn't X actually dangerous?", you will most likely get a confirming answer — even when the real evidence is more nuanced. The model optimizes for meeting your expectation, not for contradicting you. So frame your questions as neutrally and openly as possible ("What does the research say about X?" rather than "X is harmful, right?"). For important topics, it also helps to explicitly ask for counterarguments — otherwise you only get the version that fits your framing.
!What persists despite all progress
Hallucinations, lack of true intelligence, the black-box nature, bias, and the mirroring of leading questions are not technical teething problems that will disappear with the next update. They are consequences of how LLMs fundamentally work. As long as LLMs are based on statistical prediction, these limitations will persist in some form.
Why do LLMs still hallucinate even in 2026?
Apply
What does this mean for you in practice?
-
Always verify facts. When an LLM gives you a statistic, a quote, or a source, check it yourself. Especially for important decisions — medical, legal, financial — an LLM is a starting point, not the final word.
-
Use long contexts. You can now upload entire documents and ask questions about them. Take advantage of this! Instead of summarizing information for the LLM, give it the original document and let it work with the full text.
-
Use memory intentionally. If you work regularly with an LLM, configure Memory or Projects so the model knows your context. But check occasionally what has been stored — the model does not always remember things correctly.
-
Stay critical on sensitive topics. On topics like health, law, finance, or culture, model biases can lead to problematic answers. Use LLMs as tools here, not as authorities.
Reflect
The only constant in the world of LLMs is change. What was considered a fundamental limitation in 2023 — no internet access, tiny context windows, no memory — has long been resolved by 2026. The pace of this development is unprecedented in the history of technology.
At the same time, there are limitations that run deeper. Hallucinations, lack of true intelligence, and bias are not bugs that can simply be fixed. They are properties of the technology itself. The most important skill when working with LLMs is therefore not blind trust, but informed use: knowing what an LLM can do, and knowing where you are better off thinking for yourself.
You are using an LLM to review an important contract. The LLM says: 'Section 12 of the Civil Code permits this.' What do you do?