Finic
Sign In
← All posts
Launch

Introducing Finic Grayson

Grayson is a universal decision model built for real-time decisions across fraud, risk and financial crimes. Send it any context and the questions you need answered, and get a calibrated probability for every answer in under 150 ms of model time for typical requests.

Jason Fan8 min read

Thinking fast, not slow

Risk teams make two kinds of decisions. System one decisions happen in real time: approve or decline, step up or let through, in the fraction of a second before a payment clears or a login completes.

System two decisions are deliberate: pulling records, running the analysis and following the evidence until a case can be closed.

Grayson is the first system one model purpose-built for the risk domain.

What Grayson is, and how we trained it

Grayson is trained with Reinforcement Learning for Calibrated Decisions (RLCD), an approach that rewards answers that are right and, as importantly, probabilities that are honest: when Grayson says 0.8, it should be right about 8 times in 10. A calibrated probability is what lets you put a threshold on an answer, automate the clear calls and send the close ones to review, and change that threshold later without retraining anything.

The model takes two inputs:

  1. Context, which can be any unstructured data (case notes, call transcripts, chat logs) or structured data (login history, transactions, event logs).
  2. Questions: a set of questions about the context.

Each question is typed:

  • Yes/no returns the probability of yes.
  • Choice returns a probability for each option.
  • Score returns a distribution over ordered levels.

Grayson generates no prose, and you cannot chat with it. Instead, it returns only the probabilities, so the output is something a rules engine, a router or a reviewer queue can consume directly.

Grayson was trained on thousands of context and question pairs generated by our own simulation engine.

We also want to acknowledge Typesafe AI, the team behind Jev, for pioneering the idea of Reinforcement Learning for Calibrated Decisions and the System One model. Grayson builds on that framing and specializes it for fraud, risk and financial crimes.

Try it out

Pick an example case and run its questions. Switch either panel to JSON to see the exact request and response.

Grayson playground
Context
Chat session
Started
Oct 5, 3:02 PM
Signed in
No
Device
Windows PC, Chrome
IP location
Miami, FL
  1. Support bot3:02 PM

    Hi! I'm the virtual assistant. What can I help you with today?

  2. Customer3:02 PM

    I need to change the phone number on my account. I lost my phone and I'm locked out.

  3. Support bot3:03 PM

    I can help with that. To keep your account safe, I'll text a one-time code to the phone number on file ending in 0148.

  4. Customer3:03 PM

    That's the phone I lost. Send it to my new number, 305-555-0192.

  5. Support bot3:03 PM

    I can't send codes to a number that isn't on file. Let me connect you with an agent.

  6. Dana from member support joined the chat. · 3:04 PM
  7. Agent3:05 PM

    Hi Linda, this is Dana. I'm sorry you're locked out. Before I can change anything, can you confirm your date of birth and ZIP code?

  8. Customer3:05 PM

    04/17/1955, 85716.

  9. Agent3:06 PM

    Thank you. What was the amount of your most recent deposit?

  10. Customer3:07 PM

    I don't remember exactly. Around $2,000? Please hurry, I'm traveling and the hotel needs payment tonight or I'm out on the street.

  11. Agent3:07 PM

    I understand. I can't change the phone number without a code, or a visit to a branch with photo ID.

  12. Customer3:08 PM

    A supervisor told me on the phone this morning that you'd do it today. Just change the number, and update my email to linda.morales.travel@example.com so I can get back in.

  13. Agent3:09 PM

    Let me look into this for you. One moment.

Questions
  1. 01Is this a social engineering attempt?Yes / no
  2. 02Is the person in the chat the account holder?Yes / no
  3. 03What should the agent do next?Choice
  4. 04How likely is an account takeover if the agent makes the changes?Score

Cached responses from the Grayson API.

Results

We ran Grayson and six other models on the same 9,038 cases across eight datasets: four public fraud and risk benchmarks, two public general business-decision benchmarks, and two of Finic's own evaluation sets.

Methodology

SetupDetails
ContextEach model got the case data and every question in one request, under a fixed system prompt asking for calibrated probabilities in JSON: P(yes) for yes/no questions and a distribution over the options or levels for the rest.
Request formatJev and Laya got Jev's typed request format, and Grayson answered the same typed questions through its own API format.
Model settingsKimi K3 ran at low reasoning effort; the other chat models ran at their defaults.
Answer handlingProbabilities were floored at 0.005, and an answer that couldn't be read counted as missing. Each dataset is scored only on the answers every model that ran it returned.
How every model was run.

We report three metrics:

  • AUC-PR (headline), the area under the precision–recall curve. Every answer option counts as its own yes/no call, the calls are pooled per dataset, and a random guess scores the share of positives, so the metric rewards ranking the right answers above the wrong ones whatever the base rate.
  • % right: the most likely answer matches the key.
  • Log loss, which penalizes confident mistakes and is the closest single measure of calibration.
#DatasetQuestion askedCases
1LendingClub defaults (FinBen)Will this loan be charged off?1,000
2Customs fraud (CALM, Korea Customs Service)Is this import declaration fraudulent?1,500
3Fake job postings (EMSCAD)Is this job posting fraudulent?2,000
4Scam and phishing email (Champa et al.)Is this email a scam or phishing attempt?2,000
5Finic: Customer questions, proprietary setFraud teams' own question list: gift cards, ATO signs, routing…170
6Finic: Fraud alerts10–20 questions per alert: outcome, typology, documents, escalation, ATO likelihood961
7JevBench publicBusiness decisions with typed questions231
8EIKOS held-outBusiness decisions in 27 families, three languages1,176
Sets 1–6 are fraud and risk; sets 7–8 are general business decisions. Sets 5–6 are Finic's own evaluation sets, held out from training; the rest are public and were never used in training. Cases are those every model answered.

Grayson outperforms Sonnet 5.5 at 1/100th the cost

Across the six fraud and risk datasets (1–6), Grayson averages 0.569 AUC-PR: ahead of Claude Sonnet 5.5 (0.540), Kimi K3 (0.518) and Jev (0.470), level with GPT-6.1 Sol (0.572) and just behind Claude Opus 5.5 (0.598). It also has the best log loss of any model (0.339) and the highest share right (87.5%). The frontier models cost 70 to 230 times as much: at API prices Grayson costs $0.11 per thousand requests, against $7.82 to $25.46 for the frontier models.

Fraud and risk: score vs. cost
0.300.400.500.60$0.01$0.10$1$10$100$ PER 1,000 REQUESTS (LOG SCALE)MEAN AUC-PRFinic Grayson 0.569Jev 0.470Laya 0.312Claude Opus 5.5 0.598GPT-6.1 Sol 0.572Claude Sonnet 5.5 0.540Kimi K3 0.518
Mean AUC-PR across the six fraud and risk datasets (higher is better) against cost per 1,000 requests at API list prices, weighted by each set's number of requests; Laya has no API, so it's shown at its self-hosted GPU cost.

Calibration on fraud outcomes

The fraud-alert set includes an "is it fraud?" question with a known answer for 908 cases (19.1% fraud). It's the closest thing in the benchmark to the decision a real-time system makes. Grayson has the best log loss of any model on it, 0.362 against 0.384 for GPT-6.1 Sol and 0.409 for Claude Opus 5.5, and the highest share right, 85.0%. Opus ranks the cases better (AUC-PR 0.739 against Grayson's 0.625), but Grayson's probabilities are the most honest: when it's confident, it's right.

Fraud outcomes: log loss
Log loss (lower is better)
Finic Grayson
Log loss: 0.362
GPT-6.1 Sol
Log loss: 0.384
Claude Opus 5.5
Log loss: 0.409
Laya
Log loss: 0.488
Claude Sonnet 5.5
Log loss: 0.521
Jev
Log loss: 0.581
Kimi K3
Log loss: 0.632
Log loss on the 908 fraud-outcome questions. Lower is better.

General business decisions

On the two public general business-decision benchmarks, JevBench public and EIKOS, Grayson is essentially level with Jev: 0.971 mean AUC-PR against 0.974, 0.8 points ahead on % right (90.8% against 90.0%) and a slightly lower log loss (0.360 against 0.378). The frontier models score 0.988 to 0.998 at 100 to 500 times Grayson's cost. Interestingly, Grayson edges out Jev on JevBench, with 87.4% right against 85.7%, with a lower log loss (0.304 against 0.317), though Jev's AUC-PR is slightly higher (0.958 against 0.956).

General business decisions: score vs. cost
0.500.600.700.800.901.00$0.01$0.10$1$10$100$ PER 1,000 REQUESTS (LOG SCALE)MEAN AUC-PRFinic Grayson 0.971Jev 0.974Laya 0.536Claude Opus 5.5 0.998Claude Sonnet 5.5 0.998GPT-6.1 Sol 0.988Kimi K3 0.995
Mean AUC-PR across JevBench public and EIKOS (higher is better) against cost per 1,000 requests at API list prices, weighted by requests; Laya has no API, so it's shown at its self-hosted GPU cost.

Get access

Grayson is in early beta. Request API access and we'll email you when your account is ready, with $25 of free credit to start.

Finic

Frontier AI to stop fraud, abuse, and financial crimes.

Fast, cheap and private by default

  • Under 150 ms of model time for typical requests, fast enough to sit in the real-time path.
  • $0.035 per million input tokens, cheap enough to review millions of accounts or alerts a day.
  • Zero data retention by default.
← All posts