A practical guide for market participants using ChatGPT, Claude, and AI coding assistants like Claude Code, Antigravity, Cursor, and Codex in their daily workflow
You ask an AI assistant for a list of research papers by a well-known quant researcher. It gives you five titles, complete with co-authors and publication years. You feel productive. You paste the citations into your research note.
Except none of those papers exist.

Or here is a version closer to home. You ask Claude Code to write a backtest for a gap-and-go strategy on Nifty 50 stocks. It produces clean, readable Python. The logic looks right. You run it and the equity curve is beautiful. You scale up position sizing and take it live. Two weeks later you realize the code was using tomorrow’s open price to decide today’s entry. The backtest was silently peeking into the future.
This is what hallucination looks like in a trading workflow. It is not a model crashing or refusing to answer. It is a confident, articulate, often persuasive wrong answer that looks exactly like a correct one. For traders and investors, the tax on missing it is paid in capital.
What exactly is an AI hallucination?
Answer: An AI hallucination is when a language model generates information that sounds authoritative and plausible but is factually false, fabricated, or unsupported by real data.
Anthropic, the company behind Claude, has been public about the fact that while their models hallucinate far less than a year ago, the problem is not solved for any frontier AI system. Hallucinations show up as fake citations, invented statistics, wrong facts about real people and real events, and in the coding context, logic that looks correct but does something different from what it claims to do. The danger is not the error itself but the tone of confidence wrapped around it.
Think of it the way one AI educator describes it: the LLM is that arrogant friend who refuses to say “I don’t know.” Ask them anything, and they will give you an answer. Most of the time they are a tremendous resource because they genuinely know a lot. Occasionally they invent something to avoid looking unsure, and they deliver the invention with the exact same tone they use for everything else. You cannot tell the two apart from tone alone.
Why do AI systems hallucinate in the first place?
Answer: AI models predict the most likely next words based on patterns in training data, so when information is thin, missing, or past the training cutoff, they generate plausible-sounding content instead of admitting uncertainty.
Two reasons matter most for traders.
The training cutoff. Every model has a date beyond which it has no knowledge. If you ask ChatGPT or Claude about an earnings release from last week, a SEBI circular from yesterday, or a corporate action announced this morning, the model was not trained on that information. It should say so. Often it does. Sometimes it does not, and instead it produces a confident answer based on what similar events usually look like.
Thin training signal. Large language models learn by reading enormous volumes of text and becoming extraordinarily good at predicting what token comes next. When you ask about something well-documented, such as the Black-Scholes formula or the 2008 financial crisis, the model has seen that material hundreds of thousands of times and reproduces it reliably. When you ask about a small-cap Indian pharma stock, the intraday microstructure of a specific option strike, or a niche regulatory provision, the signal gets weak. The model still tries to be helpful. It pattern-matches to what similar answers usually look like, and produces something that sounds right.
There is also a third, subtler failure mode specific to coding assistants. Tools like Claude Code, Antigravity, Cursor, and Codex generate code that compiles and runs. They hallucinate at the level of logic, not syntax. The code works. It just does not do what you think it does.
Why should traders and investors care more than most users?
Answer: In markets, a single confidently wrong fact about earnings, regulation, or logic in a backtest can translate directly into real financial loss, making hallucination risk more consequential than in casual use.
A student using AI to draft a history essay pays a grade penalty if the AI invents a citation. A trader who acts on a hallucinated earnings number, a fabricated circular, or a backtest with subtle look-ahead bias loses real capital. The asymmetry matters.
A few scenarios where traders are especially exposed.
Research summaries where the AI paraphrases a report that does not actually contain the claim being attributed to it. Regulatory questions where the AI describes a rule that sounds plausible but was repealed, never enacted, or applies to a different jurisdiction. Fundamental analysis where the AI cites specific revenue, margin, or balance sheet numbers that are close to real but meaningfully off. News-adjacent queries about recent events or corporate actions where the AI confidently reports something that never happened because the training data did not cover the timeframe.
And then there is coding, which deserves its own section.
The specific problem with AI coding assistants for strategy development
Answer: AI coding tools produce code that runs cleanly but may contain subtle logic errors in data handling, look-ahead bias, or assumptions about market mechanics that only surface when real money is at risk.
If you use Claude Code, Antigravity, Cursor, or Codex to build trading strategies, you are in a category of hallucination risk that most AI users never encounter. The failure mode is specific and worth naming.
Look-ahead bias in backtests. The model writes a loop that uses today’s close to decide today’s entry, or shifts the wrong column when aligning signals with prices. The code runs, the equity curve looks great, and the strategy is unreproducible in live trading. This is the most expensive single hallucination category I have seen.
Data handling assumptions. The model assumes your data is split-adjusted when it is not, or vice versa. It assumes dividend handling you never specified. It resamples intraday data to daily using the last tick rather than the close, or mixes timezones silently between the exchange and your system clock.
Survivorship bias. The model pulls a current Nifty 500 constituent list and backtests over ten years, not accounting for the fact that those constituents were not in the index ten years ago. Performance looks great because you are only trading stocks that survived.
Broker API misuse. The model invents a parameter name, uses a deprecated endpoint, or hallucinates an order type that does not exist on the broker you are connecting to. This is especially common when working with Indian broker APIs where the training signal is thinner than for US brokers.
Option Greeks and pricing formulas. The model writes a Black-Scholes implementation that is almost right. The decimal on volatility is wrong, the risk-free rate is annualized incorrectly, or the time to expiry is in trading days rather than calendar days. Your position sizing based on delta is silently off.
Indicator calculations. The model writes an RSI, a VWAP, or a supertrend implementation that differs subtly from the standard definition. Your signals fire at different times from your TradingView or Amibroker reference. You do not notice until you compare side by side.
The unifying theme is that none of these errors crash the program. They just give you wrong PnL, wrong signals, or wrong fills. You find them only if you look for them.
Where are hallucinations most likely to occur in a trading workflow?
Answer: Hallucinations cluster around specific numbers, obscure tickers, recent events, niche regulations, citations, and code involving market mechanics, which unfortunately overlaps heavily with what traders actually ask AI to do.
Map the known risk zones onto a working trader’s day.
Specific numbers are everywhere in trading: EPS estimates, option Greeks, volatility surfaces, margin requirements. Obscure topics include low-float small caps, sector-specific jargon, and emerging-market microstructure. Recent events cover anything after the training cutoff. Citations are the foundation of any serious investment thesis. Real but not widely known entities describe most of the Indian mid and small cap universe. Code involving market mechanics is where AI coding assistants concentrate their subtlest errors.
In other words, the exact set of queries a trader is most likely to send to an AI is also the exact set most vulnerable to hallucination.
How can you catch hallucinations when they happen?
Answer: Ask for sources and verify them, challenge the model’s confidence, start fresh chats to cross-check, and for code, test against known cases before trusting any logic with real capital.
Here is a layered defense that works in practice.
Ask for sources and verify them. When the AI cites a paper, a regulation, or an article, do not stop at the citation existing. Open it and confirm the cited source actually makes the claim being attributed to it. Hallucinated sources sometimes point to real documents that do not contain the claim, which is a subtler and more dangerous failure than a fully invented citation.
Give the AI permission to be uncertain. Tell the model upfront that it is okay to say “I don’t know.” This small framing change measurably reduces confabulation because the model no longer feels pressure to produce an answer-shaped response at all costs.
Probe confidence directly. Ask the AI how confident it is in a specific claim, and whether any part of its answer might be wrong. Often the model has internal uncertainty it did not surface in the original response. A follow-up question like “which of these five claims are you least confident about” can reveal exactly where to focus verification.
Use a fresh chat as a second opinion. Paste the AI’s answer into a new conversation and ask it to find errors, check the reasoning, and confirm that the claims hold up. A fresh context window often catches problems that the original conversation would defend.
Cross-reference against primary sources for anything that matters. For regulatory questions, go to the SEBI, RBI, or exchange circular directly. For company fundamentals, go to the annual report, the BSE or NSE filings, or the earnings presentation. For price and volume data, use your broker feed or a verified market data provider. The AI is a drafting partner, not a source of truth.
Be specifically skeptical of exact numbers, dates, and names. These are the highest-risk items because they are precise enough to be checkable and consequential enough to matter. If an answer hinges on a number, verify the number.
A separate verification discipline for AI-generated strategy code
Answer: Treat every piece of AI-generated trading code as guilty until proven innocent through deliberate testing against known cases, reference implementations, and paper trading before any real capital is committed.
When you use Claude Code, Antigravity, Cursor, or Codex to build a strategy, verification is not optional. Build these steps into your workflow permanently.
Test against a known case. Before trusting any AI-generated indicator, run it on a dataset where you already know the correct output from Amibroker, TradingView, or Pine Script. If your AI-generated RSI does not match the reference value bar for bar, something is wrong. Do this for every indicator, every time, before using it in a strategy.
Check for look-ahead bias explicitly. After the AI writes a backtest, ask a fresh chat to review the code specifically for look-ahead bias, future leakage, and survivorship bias. Be explicit that the reviewer should assume the code is broken and find the break. A clean bill of health from a skeptical second pass is more meaningful than passing your own first read.
Paper trade before you fund. No matter how clean the backtest looks, run the strategy in paper mode or with minimum capital for a period that covers at least one volatility regime. Reconcile every fill, every signal, and every PnL number against what you expected. Mismatches are almost always bugs, not bad luck.
Reconcile broker API behavior against documentation. If the AI generates code that places orders, verify every parameter against the broker’s official API documentation. AI assistants hallucinate parameter names, order types, and exchange codes more often than any other single category when working with Indian broker APIs.
Version control and human review. Every AI-generated change that touches live trading logic gets committed to git with a clear message and reviewed as if a junior developer wrote it. If the change is too complex for you to review line by line, it is too complex to deploy.
Keep humans in the loop on anything consequential. AI can draft the strategy, write the code, suggest the parameters, and explain the logic. The decision to commit real capital is yours alone, made after verification you performed yourself.
Is this going to get better?
Answer: Yes, hallucination rates are falling meaningfully with each model generation, but the problem is unlikely to reach zero soon, and trader-facing use cases will remain sensitive for the foreseeable future.
The trajectory is genuinely encouraging. Anthropic’s public stance is that they measure hallucination rates, track them over time, and see consistent improvements across releases. Other labs report similar progress. Coding assistants in particular have improved substantially in the last year, both in raw capability and in self-awareness about uncertainty.
But even as the rate drops, the stakes in trading keep the required verification effort high. A 1 percent hallucination rate on a casual question is fine. A 1 percent hallucination rate on 100 regulatory claims or 100 lines of trading logic you acted on over a year is a meaningful risk.
The mental model to hold is that AI is becoming a progressively better research assistant and coding partner, not a replacement for primary sources or human review. The traders and investors who will extract the most value are the ones who build verification into their workflow as a permanent habit, treat AI output as a high-quality first draft rather than a final answer, and stay calibrated about which kinds of questions are reliably answered and which kinds still require independent confirmation.
The bottom line for market participants
AI assistants and AI coding tools are genuinely useful for traders and investors. They accelerate research, improve writing, generate strategy code, structure analysis, and make it easier to explore unfamiliar topics. That value is real and it is growing.
Hallucinations are the tax you pay for that capability, and the tax is collected disproportionately from users who do not verify. The good news is that catching hallucinations is a learnable skill. Ask for sources. Check the sources. Give the model permission to hedge. Probe confidence. Start fresh chats to audit. Cross-reference anything that matters against a primary source. For code, test against known cases, check explicitly for look-ahead bias, reconcile against broker documentation, and paper trade before you fund.
Treat every AI output the way a disciplined investor treats a sell-side research note. Read it seriously, take the framing and the analysis on board, and then independently verify every number, every claim, and every line of code that would move your decision. That is not paranoia. That is just good research hygiene applied to a new class of tool.
The traders who will do best with AI over the next few years are not the ones who trust it the most. They are the ones who use it the most while trusting it the least.