Sentinel enhances fraud detection with graph analysis and AI-generated summaries

Vansh Dobhal’s Sentinel project leverages graph analysis, policy rules, and language generation to streamline bank card fraud investigations, providing more accurate, auditable, and real-time case summaries.

Every month, bank card-fraud teams are left to work through a small but stubborn queue of cases that still need human judgement: what happened, whether the charge was fraudulent, how far the activity spread, and whether a suspicious activity report should be filed. In a submission to Hacker House Goa and TigerGraph, Vansh Dobhal presented Sentinel, an agentic fraud investigator built to answer those questions by combining graph analysis, a deterministic policy layer and language generation for prose only. The project is designed so that the machine learning model and graph queries supply the evidence, while the language model merely turns that evidence into a readable case summary.

The structure reflects a broader shift in fraud tooling towards relationship-based analysis. TigerGraph positions its platform as a real-time graph system for AI and analytics, with particular emphasis on fraud detection, anti-money laundering and other settings where relationships between accounts, devices, people and transactions matter more than single-row signals. Its fraud-detection materials make the same case more explicitly, arguing that graph analytics and machine learning together can expose collusion, synthetic identities and other schemes that conventional systems struggle to spot. Sentinel is built around that idea, treating the graph as the place where the story of a case is actually visible.

Dobhal says the first instinct was to let a rules engine read the alert row and decide the next step. That approach quickly ran into a familiar problem: the alert itself was already the result of upstream risk scoring, so any obvious feature on the flagged transaction risked being double-counted. The project therefore shifted to a model trained on fraud versus cleared cases rather than fraud versus everything else. In the write-up, the resulting classifier reached a five-fold cross-validated AUC of 0.9465 with a Brier score of 0.0927, and its coefficients exposed some counter-intuitive patterns, including evidence that a new device flag was not, on its own, a strong fraud signal in this alert population.

The graph layer is where Sentinel tries to move beyond ordinary transaction monitoring. Dobhal describes the use of TigerVector’s cosine index over embedded case notes, installed GSQL queries and a custom weakly connected-component routine over a card-device projection to surface shared-device patterns. The project also uses asynchronous query dispatch to reduce wall-clock time, and it relies on a TigerGraph multi-hop view of the data to answer questions such as which cases resemble the current alert, which devices connect otherwise separate cards and whether a cluster looks like a narrow fraud ring or simply a broad, noisy network.

That distinction matters because the project treats some graph signals as context rather than direct evidence. The shared-device component can become enormous even at conservative degree limits, which means a large connected set does not automatically indicate a ring. Sentinel therefore separates blast-radius information from the actual decision signal, using narrower filters to decide whether a case should be treated as suspicious. Dobhal says the system’s ledger of actions comes from a policy engine that encodes the bank’s own rules, while the large language model is limited to writing the analyst summary and the SAR narrative.

The project also places tight controls around that generated text. Dobhal says each prompt is restricted to an allow-list of case identifiers and the output is checked afterwards for any invented references. The policy engine, meanwhile, enforces a set of invariants so that the final answer file stays internally consistent: if a report is filed, the action list must say so; if connected cards are identified, the system must flag the shared element that caused the connection; and the verdict must align with the model probability. The point is not simply to make the agent more accurate, but to make it auditable.

In the test run described in the article, Sentinel produced twelve fraud and eight legitimate outcomes, with four cases triggering SARs. One highlighted example involved a shared device profile that connected multiple cards and attached prior confirmed fraud to the same device, causing the system to escalate to report filing and monitoring actions. Another case, by contrast, was treated as a recurring legitimate charge rather than a card compromise, showing how the system can avoid unnecessary card blocking when the evidence points to a pattern the customer would reasonably recognise. Dobhal says the broader lesson is that customer response can settle a borderline case, and that human-in-the-loop logic has to remain part of the design.

The broader aim of the project is not to replace analysts, but to compress the time and ambiguity of their work into a traceable workflow. That aligns with the wider positioning of TigerGraph and with other Sentinel-branded platforms that stress real-time investigation, human oversight and compliance-friendly audit trails. In this version, the promise is narrower and more technical: fraud decisions should be grounded in graph evidence, policy should be deterministic, and any prose generated by an agent should be there to explain the conclusion rather than create it.

Disclaimer: This article is intended to inform and educate, not to recommend or endorse any financial product, investment or strategy. Please consider your own financial circumstances and seek professional advice where appropriate before making financial decisions.