Coldfire
ProductUse casesEarly access
Get access
Synthetic data for AI teams, now in early access

The training data you need, without the data you can't use.

Coldfire turns a schema and a handful of examples into large, balanced datasets for fine-tuning, evaluation and testing, and checks every row before it reaches your model.

Request early accessSee how it works
support-intents/run 14
1,284 of 2,000 rows

Schema

text
string
label
enum(6)
channel
email | chat

Seeds

12 hand-written examples

Rules

  • Vary tone and length
  • Mix typos and slang
  • No real names or emails
  • Hold out eval_v3
  1. 1279

    Two charges hit my card for the same order. Can you reverse one?

    billing_disputeschema, dedupe, PII, eval overlap passed
  2. 1280

    app keeps logging me out after the update?? android 14

    bug_reportschema, dedupe, PII, eval overlap passed
  3. 1281

    Where do I change the email that receipts go to?

    accountschema, dedupe, PII, eval overlap passed
  4. 1282

    Where can I update the email my receipts are sent to?

    accountrejected: near-duplicate of 1281
  5. 1283

    Cancel at the end of this cycle please, not today.

    cancellationschema, dedupe, PII, eval overlap passed
  6. 1284

    Is there a nonprofit discount, and does it stack with annual?

    pricingschema, dedupe, PII, eval overlap passed

Label balance

  • billing_dispute19%
  • account17%
  • bug_report18%
  • cancellation15%
  • pricing16%
  • other15%

Validators

Near-duplicates removed
3.1%
PII matches
0
Overlap with eval_v3
0.0%
Illustrative example, not customer data.
Product

Built for data you can actually train on.

Generating text is easy. Generating a dataset that's varied, balanced, clean and separate from your evals is the hard part. That's what Coldfire is for.

Coverage, not repetition

Coldfire maps the space your seeds describe and steers generation toward the combinations you haven't covered yet: rare labels, odd phrasings, edge cases.

Your seedsGeneratedStill to cover

Every row is checked

Rows that fail a validator never make it into the export. You see what was dropped and why.

  • Schema and label setpass
  • Near-duplicate detection62 dropped
  • PII and secrets scanpass
  • Your custom rules4 dropped

Evals stay clean

Point Coldfire at your held-out set and it keeps generated training rows from overlapping with it, so benchmark gains are real.

train
eval_v3

overlap 0.0%

Export with provenance

JSONL, CSV, Parquet or a Hugging Face dataset. Each row remembers its seed, settings and validator results, so you can audit or regenerate any slice.

JSONLParquetCSVHF Hub
{"text": "Two charges hit my card for the same order.",
 "label": "billing_dispute",
 "_meta": {"seed": "s07", "run": 14,
           "checks": ["schema", "dedupe", "pii", "eval_overlap"]}}
Use cases

One generator, four jobs.

For ML engineers and AI product teams who need more data than they can collect, label or legally use.

Fine-tuning sets

Instruction and preference pairs for the tasks your base model is weak at, in the format your trainer expects.

12 seeds → 8,000 instruction pairs

Evals and red-team suites

Held-out test sets and adversarial prompts that stay separate from training data, so scores mean something.

jailbreak, refusal and edge-case suites

Tabular data for ML

Rows with the same shape and correlations as data you can't share, for prototyping and model development.

schema in, Parquet out

Fixtures for AI features

Repeatable inputs for the prompts, agents and pipelines you ship, so regressions show up before users do.

versioned with every run

Early access

We're onboarding a small group of teams.

Early access is hands-on. You get a working dataset quickly, and we learn what to build next from real projects.

  1. 1

    Tell us what you're building

    Send a short note: the model or feature, the data you have, and the data you wish you had.

  2. 2

    Working session on your schema

    We generate a first dataset together from your schema and a few seed examples, and review the rejects.

  3. 3

    Run it on a real project

    Use Coldfire on one fine-tune or eval suite, with direct access to the team building it.

Your seeds stay yours

Seed data and generated datasets are never used to train shared models.

No production data needed

Start from a schema and hand-written examples. Nothing sensitive has to leave your systems.

Delete on request

Anything you share during early access is removed whenever you ask.

The details are in our privacy policy.

Stop waiting on data. Generate it.

Tell us what you're training or evaluating, and we'll set up a working session on your own schema.

Request early accesshello@coldfire.ai
Coldfire© 2026
PrivacyTermshello@coldfire.ai