Fine-tuning sets
Instruction and preference pairs for the tasks your base model is weak at, in the format your trainer expects.
12 seeds → 8,000 instruction pairs
Coldfire turns a schema and a handful of examples into large, balanced datasets for fine-tuning, evaluation and testing, and checks every row before it reaches your model.
Schema
Seeds
12 hand-written examples
Rules
Two charges hit my card for the same order. Can you reverse one?
app keeps logging me out after the update?? android 14
Where do I change the email that receipts go to?
Where can I update the email my receipts are sent to?
Cancel at the end of this cycle please, not today.
Is there a nonprofit discount, and does it stack with annual?
Label balance
Validators
Generating text is easy. Generating a dataset that's varied, balanced, clean and separate from your evals is the hard part. That's what Coldfire is for.
Coldfire maps the space your seeds describe and steers generation toward the combinations you haven't covered yet: rare labels, odd phrasings, edge cases.
Rows that fail a validator never make it into the export. You see what was dropped and why.
Point Coldfire at your held-out set and it keeps generated training rows from overlapping with it, so benchmark gains are real.
overlap 0.0%
JSONL, CSV, Parquet or a Hugging Face dataset. Each row remembers its seed, settings and validator results, so you can audit or regenerate any slice.
{"text": "Two charges hit my card for the same order.",
"label": "billing_dispute",
"_meta": {"seed": "s07", "run": 14,
"checks": ["schema", "dedupe", "pii", "eval_overlap"]}}For ML engineers and AI product teams who need more data than they can collect, label or legally use.
Instruction and preference pairs for the tasks your base model is weak at, in the format your trainer expects.
12 seeds → 8,000 instruction pairs
Held-out test sets and adversarial prompts that stay separate from training data, so scores mean something.
jailbreak, refusal and edge-case suites
Rows with the same shape and correlations as data you can't share, for prototyping and model development.
schema in, Parquet out
Repeatable inputs for the prompts, agents and pipelines you ship, so regressions show up before users do.
versioned with every run
Early access is hands-on. You get a working dataset quickly, and we learn what to build next from real projects.
Send a short note: the model or feature, the data you have, and the data you wish you had.
We generate a first dataset together from your schema and a few seed examples, and review the rejects.
Use Coldfire on one fine-tune or eval suite, with direct access to the team building it.
Seed data and generated datasets are never used to train shared models.
Start from a schema and hand-written examples. Nothing sensitive has to leave your systems.
Anything you share during early access is removed whenever you ask.
The details are in our privacy policy.
Tell us what you're training or evaluating, and we'll set up a working session on your own schema.