← All posts
Product

Which AI model works best in Groundtruth?

Small models for quick questions, mid-size for setup, the best models for contracts and writing. One pick from Anthropic, OpenAI and Google for every stage, with what each costs.

You can now bring your own key from Anthropic (Claude), OpenAI or Google (Gemini). Each offers a small, a mid-size and a top model, and the price gap between them is big: up to 100 times per question. So which one should you pick?

It depends on what you expect from it. The rule of thumb:

  • "The answer is in our pages, just find it quickly": a small model. It's fast and costs almost nothing.
  • "Read something messy and turn it into structure": a mid-size model. It follows instructions closely and costs a fraction of the best.
  • "A mistake costs us money" or "it has to read well": a top model. You use it less often, so the higher price matters less.

Below is how that works out at each stage in Groundtruth, with one option from every supplier.

Stage 1: Setting up from your website

What you expect: read five pages of marketing copy, pull out what's true about your company, and write it up as tidy pages, FAQ answers and records. Don't invent anything.

This is structured extraction from messy input: the sweet spot of a mid-size model. Small models tend to miss details or bend the format; top models do it well but cost more than they add, for a one-time job.

Supplier Pick About
Anthropic Claude Sonnet 5.5 $0.06 per import
OpenAI GPT-6.1 Sol $0.06 per import
Google Gemini 3.8 Flash $0.02 per import

Stage 2: Everyday questions

What you expect: "What's our refund policy?" answered in a few sentences, with citations, in a second.

Groundtruth does the hard part before the model sees anything: it finds the five best-matching passages and sends only those. The model's job is to read a page's worth of text and answer from it. Small models are very good at exactly this, and they're the fastest. It's also why the free credit you get when you sign up runs on Claude Haiku 5.5.

Supplier Pick Per question Per 1,000 questions
Anthropic Claude Haiku 5.5 $0.0004 $0.40
OpenAI GPT-6 Luna $0.0004 $0.40
Google Gemini 3.5 Flash-Lite $0.0015 $1.50

Move up a tier if your questions need reasoning across sources ("Do our travel and expense policies contradict each other?") rather than finding a fact.

Stage 3: Contracts

What you expect: read a 15-page supplier contract and get the end date, the notice period and the renewal terms exactly right, because the Contract Agent turns them into a cancel-by date.

One wrong date can mean another year of a contract you meant to cancel. Contracts don't come in often, so paying more per contract is cheap insurance. Use a top model. You still check every field before it's saved.

Supplier Pick Per contract
Anthropic Claude Opus 5.5 $0.06
OpenAI GPT-6 Astra $0.15
Google Gemini 3.1 Pro (preview) $0.03

Stage 4: Writing policies in Write mode

What you expect: alternatives that sound like you, Lab marks that point at the sentence you'd have flagged yourself, and trims that cut filler, not meaning.

This is judgement and taste, where top models clearly pull ahead. Smaller models suggest safe, generic wording and mark too much or too little. For policy writers who spend hours in a document, the difference is worth a few cents.

Supplier Pick Per Lab run on 3,000 words
Anthropic Claude Opus 5.5 $0.05
OpenAI GPT-6 Astra $0.12
Google Gemini 3.1 Pro (preview) $0.03

Stage 5: Your assistant, connected

What you expect: ask Claude, ChatGPT or Gemini about your company, in the chat you already use.

Here you don't need a key at all. When you connect your assistant, Groundtruth sends it the cited sources and your assistant writes the answer with the model you already chose in its own app, on your own plan. Nothing is billed by Groundtruth.

If you can only pick one

Today a workspace uses one model for everything in the app, so pick for what your team does most:

Your team mostly… Anthropic OpenAI Google
…asks questions Claude Haiku 5.5 GPT-6 Luna Gemini 3.5 Flash-Lite
…does a bit of everything (most teams) Claude Sonnet 5.5 GPT-6.1 Sol Gemini 3.8 Flash
…works on contracts or writes policy Claude Opus 5.5 GPT-6 Astra Gemini 3.1 Pro (preview)

Not sure? Start with the middle row. It answers everyday questions well, handles setup comfortably and is good enough for most writing. For 1,000 questions a month that's about $8 on Claude Sonnet 5.5 or GPT-6.1 Sol, and $3 on Gemini 3.8 Flash. You can switch models any time in Settings → AI; nothing in your workspace changes when you do.

Which supplier? If you already pay for one, use it. Claude is our default and what we test against first; OpenAI and Gemini work in the same places with the same citations and the same labels on everything AI writes.

How we estimated the costs

We used each supplier's list price on 9 October 2026 and typical sizes in Groundtruth: a question is about 2,500 tokens in and 300 out, a website import about 8,000 in and 4,000 out, a contract about 10,000 in and 1,000 out, a Lab run about 4,000 in and 1,500 out. Models that reason before answering use more output tokens than that; we ask for low effort where the model supports it. Prices change, and Google has announced that Gemini 3.8 Flash doubles in price on 1 January 2027, so treat these as orders of magnitude. Your provider bills you directly; Groundtruth adds no markup.

Ready? Add your key in Settings → AI in your workspace.