Introducing Finic Grayson
Grayson is a universal decision model built for real-time decisions across fraud, risk and financial crimes. Send it any context and the questions you need answered, and get a calibrated probability for every answer in under 150 ms of model time for typical requests.
Thinking fast, not slow
Risk teams make two kinds of decisions. System one decisions happen in real time: approve or decline, step up or let through, in the fraction of a second before a payment clears or a login completes.
System two decisions are deliberate: pulling records, running the analysis and following the evidence until a case can be closed.
Grayson is the first system one model purpose-built for the risk domain.
What Grayson is, and how we trained it
Grayson is trained with Reinforcement Learning for Calibrated Decisions (RLCD), an approach that rewards answers that are right and, as importantly, probabilities that are honest: when Grayson says 0.8, it should be right about 8 times in 10. A calibrated probability is what lets you put a threshold on an answer, automate the clear calls and send the close ones to review, and change that threshold later without retraining anything.
The model takes two inputs:
- Context, which can be any unstructured data (case notes, call transcripts, chat logs) or structured data (login history, transactions, event logs).
- Questions: a set of questions about the context.
Each question is typed:
- Yes/no returns the probability of yes.
- Choice returns a probability for each option.
- Score returns a distribution over ordered levels.
Grayson generates no prose, and you cannot chat with it. Instead, it returns only the probabilities, so the output is something a rules engine, a router or a reviewer queue can consume directly.
Grayson was trained on thousands of context and question pairs generated by our own simulation engine.
We also want to acknowledge Typesafe AI, the team behind Jev, for pioneering the idea of Reinforcement Learning for Calibrated Decisions and the System One model. Grayson builds on that framing and specializes it for fraud, risk and financial crimes.
Try it out
Pick an example case and run its questions. Switch either panel to JSON to see the exact request and response.
Cached responses from the Grayson API.
Results
We ran Grayson and six other models on the same 9,038 cases across eight datasets: four public fraud and risk benchmarks, two public general business-decision benchmarks, and two of Finic's own evaluation sets.
Methodology
| Setup | Details |
|---|---|
| Context | Each model got the case data and every question in one request, under a fixed system prompt asking for calibrated probabilities in JSON: P(yes) for yes/no questions and a distribution over the options or levels for the rest. |
| Request format | Jev and Laya got Jev's typed request format, and Grayson answered the same typed questions through its own API format. |
| Model settings | Kimi K3 ran at low reasoning effort; the other chat models ran at their defaults. |
| Answer handling | Probabilities were floored at 0.005, and an answer that couldn't be read counted as missing. Each dataset is scored only on the answers every model that ran it returned. |
We report three metrics:
- AUC-PR (headline), the area under the precision–recall curve. Every answer option counts as its own yes/no call, the calls are pooled per dataset, and a random guess scores the share of positives, so the metric rewards ranking the right answers above the wrong ones whatever the base rate.
- % right: the most likely answer matches the key.
- Log loss, which penalizes confident mistakes and is the closest single measure of calibration.
| # | Dataset | Question asked | Cases |
|---|---|---|---|
| 1 | LendingClub defaults (FinBen) | Will this loan be charged off? | 1,000 |
| 2 | Customs fraud (CALM, Korea Customs Service) | Is this import declaration fraudulent? | 1,500 |
| 3 | Fake job postings (EMSCAD) | Is this job posting fraudulent? | 2,000 |
| 4 | Scam and phishing email (Champa et al.) | Is this email a scam or phishing attempt? | 2,000 |
| 5 | Finic: Customer questions, proprietary set | Fraud teams' own question list: gift cards, ATO signs, routing… | 170 |
| 6 | Finic: Fraud alerts | 10–20 questions per alert: outcome, typology, documents, escalation, ATO likelihood | 961 |
| 7 | JevBench public | Business decisions with typed questions | 231 |
| 8 | EIKOS held-out | Business decisions in 27 families, three languages | 1,176 |
Grayson outperforms Sonnet 5.5 at 1/100th the cost
Across the six fraud and risk datasets (1–6), Grayson averages 0.569 AUC-PR: ahead of Claude Sonnet 5.5 (0.540), Kimi K3 (0.518) and Jev (0.470), level with GPT-6.1 Sol (0.572) and just behind Claude Opus 5.5 (0.598). It also has the best log loss of any model (0.339) and the highest share right (87.5%). The frontier models cost 70 to 230 times as much: at API prices Grayson costs $0.11 per thousand requests, against $7.82 to $25.46 for the frontier models.
Calibration on fraud outcomes
The fraud-alert set includes an "is it fraud?" question with a known answer for 908 cases (19.1% fraud). It's the closest thing in the benchmark to the decision a real-time system makes. Grayson has the best log loss of any model on it, 0.362 against 0.384 for GPT-6.1 Sol and 0.409 for Claude Opus 5.5, and the highest share right, 85.0%. Opus ranks the cases better (AUC-PR 0.739 against Grayson's 0.625), but Grayson's probabilities are the most honest: when it's confident, it's right.
General business decisions
On the two public general business-decision benchmarks, JevBench public and EIKOS, Grayson is essentially level with Jev: 0.971 mean AUC-PR against 0.974, 0.8 points ahead on % right (90.8% against 90.0%) and a slightly lower log loss (0.360 against 0.378). The frontier models score 0.988 to 0.998 at 100 to 500 times Grayson's cost. Interestingly, Grayson edges out Jev on JevBench, with 87.4% right against 85.7%, with a lower log loss (0.304 against 0.317), though Jev's AUC-PR is slightly higher (0.958 against 0.956).
Get access
Grayson is in early beta. Request API access and we'll email you when your account is ready, with $25 of free credit to start.