AI fraud flags: why accuracy is not enough
An AI fraud flag is a signal to investigate. It is not proof that a customer committed fraud, even when the system’s overall accuracy sounds impressive.
In iGaming, the useful question is what the tool is meant to do: flag a suspicious transaction, help answer a support enquiry or identify a different risk. The evidence needed to judge each job is different.
A small false-alarm rate can produce many alerts
Consider this fictional evaluation of 10,000 transactions. Investigators have established that 100 are fraudulent & 9,900 are legitimate. A model flags 90 of the fraudulent transactions, but also flags 99 legitimate ones.
| Actual status | Flagged | Not flagged |
|---|---|---|
| Fraudulent | 90 | 10 |
| Legitimate | 99 | 9,801 |
It catches 90 ÷ 100 = 90% of the fraud, while incorrectly flagging 99 ÷ 9,900 = 1% of legitimate transactions.
Yet only 90 ÷ 189 ≈ 47.6% of its alerts identify actual fraud. The rest are false alarms. Its overall accuracy is (90 + 9,801) ÷ 10,000 = 98.91%.
Which percentage answers your question?
Accuracy counts all correct classifications. Precision asks how many flagged cases are actually positive. Recall asks how many actual positive cases the model catches. Google’s machine-learning guidance explains why accuracy alone can mislead when one class is much rarer than the other.
Here, legitimate transactions dominate the total. That is why high overall accuracy coexists with an alert list containing more false alarms than genuine fraud.
Make the review process part of the design
Useful operational questions include who reviews an alert, what supporting evidence they can inspect, how mistakes are corrected & whether performance changes over time. The Gambling Commission’s own AI principles include appropriate human intervention, governance & assurance; they are not a statement that every operator uses the same process.
A support chatbot can explain a process or collect relevant information. That does not establish that a separate fraud classification was correct. Judge the reply’s factual accuracy & the fraud decision’s evidence separately.
Should an operator ignore an alert that might be wrong?
No. Missing genuine fraud also has costs. The practical task is to choose an appropriate response, investigate the evidence & measure both missed fraud & harm from false alarms. These example rates describe no real operator or supplier. Information only. 18+.
Sources
Sources checked 7 September 2026. Any numerical examples are illustrative.



Discussion
Ask a question, add useful context or share a source. Keep it relevant & respectful.
Comments are reviewed before publication.