afterpay-loadSynthetic Intelligence (SI)

Synthetic intelligence studio

Intelligence, engineeredfor what comes after.

afterpay-load SI designs autonomous agents, synthetic data engines and the evaluation infrastructure to trust them — built for teams who ship, not teams who demo.

Est. 2026Founded to close the gap between AI demos and AI in production
3Research pillars: agents, synthetic data, evaluation
Remote‑firstA distributed team, working async across time zones
Open notesResearch written up and shared, not kept behind closed doors

What we build

Systems built to run after launch day, not just for it.

01 / AGENTS

Autonomous agents

Persistent, tool-using agents that operate real workflows end to end, with the guardrails to fail safely instead of silently.

02 / DATA

Synthetic data engines

Generated, labelled and validated data for the cases no one could collect by hand, tuned to the distribution your system actually meets.

03 / EVALUATION

Evaluation & alignment

Benchmarks and red-teaming built to catch failure before your users do, run continuously, not once before shipping.

04 / RESEARCH

Applied research

Research that ships: every method we publish first runs in a production system, under real load, with real failure modes.

Why "after"

The first version is never the point.

Most intelligence today is built once, demoed well, and shipped.

We build what comes after: systems that keep learning,

keep checking themselves against reality, and keep getting better

long after the launch announcement is forgotten.

How we work

One pipeline, six disciplines.

Design

Scope the workflow and its failure modes before writing a line of model code.

Simulate

Build a sandbox where the agent can fail cheaply, thousands of times, before it touches anything real.

Train

Fit the system to the task, not the benchmark, using data that matches production distribution.

Evaluate

Red-team it, measure it, and refuse to ship until the failure modes are understood.

Deploy

Ship behind guardrails that fail closed, with a human path for anything the system won't decide alone.

Monitor

Watch it in the wild, catch drift early, and feed what we learn back into the next iteration.

Research notes

What we're working through right now.

In progress

Why agents fail silently

A field note on the failure modes that don't throw an error: agents that complete the wrong task confidently, and how to catch it.

AGENTS · RELIABILITY
In progress

Evaluating synthetic data without ground truth

When there is no real dataset to check against, what does "correct" even mean? Notes from building our own answer.

SYNTHETIC DATA
In progress

A minimal harness for long-horizon agents

Most agent evaluations run one step. Ours needed to run thousands. Here is the harness that made that tractable.

EVALUATION

Get in touch

Building something that needs to actually work?

Tell us about the workflow, the failure modes that worry you, and the timeline. We reply personally.

EmailUse the form, we reply from a personal address
FocusAutonomous agents, synthetic data, evaluation
AvailabilityA small number of engagements at a time, by design

Your details are used only to reply to you and are never shared.