Paul Mosquera Staff Applied AI & Platform Engineer
All writing

I Didn't Write the Code. I Wrote the Rules.

How I build new products with AI agents, told through the one project that can't afford a single shortcut: a KYC liveness platform.

  • agents
  • security
  • testing
  • method

This was the first message I sent when I started the project:

“I want to build a platform like AWS Rekognition Liveness. I want an MVP. Web first, so I can test fast. What matters is the backend logic. It should be dynamic: ask for face movements in one order, then another. And I like the idea of flashing colors on the screen to see if the light bounces off the face.

“Give me a proposal. NO CODE. NO DEMOS.”

Capital letters and all.

That last line isn’t a quirk. It’s the most important rule of how I build anything with AI agents. And on this project, a face liveness system for digital KYC, I learned why it matters more than I thought.

🔗 The full project is on GitHub: github.com/edisonpaul4/face-liveness-platform. This post is the story of how it was built.


Why KYC is the worst place to cut corners

Every digital onboarding flow has a moment of truth: a system has to decide whether the face in front of the camera is a real person, present right now, or something pretending to be one.

The attacker brings…Why a good system catches it
🖼️ A printed photoIt reacts to light, but it’s flat. No depth, no relief.
📼 A replayed videoIt can’t react to a challenge invented one second ago.
📱 A screenIt glows on its own and leaks visual artifacts.
🎭 A silicone maskIt has no pulse.

Now add AI agents to the picture. Coding agents optimize for “it works.” In most software, that’s fine. In security, “it works” is how vulnerabilities are born.

An agent that makes a test pass by bending the threat model hasn’t fixed a bug. It has opened a door.

So here’s the recipe I follow, step by step, with what actually happened on this project.

flowchart LR
    A["💡 1. Ideate<br/>no code, no demos"] --> B["🔬 2. Research<br/>concrete forks only"]
    B --> C["📜 3. Prompts in phases<br/>CLAUDE.md first"]
    C --> D["🤖 4. Delegate<br/>gate by gate"]
    D --> E["📏 5. Measure<br/>and refute"]
    E -. "what we learned<br/>goes back into the rules" .-> C

Step 1: Ideate with no code and no demos

Why ban code at the start? Because the moment code appears, the conversation shifts from what should exist to how to make this run. You stop designing and start debugging, and the architecture gets decided by whatever the first snippet happened to look like.

With code off the table, the first proposal had room to think. And the first thing it did was the most valuable thing in the entire project: it started with the attacker, not the product.

The threat model came in three levels:

  • Level 1, mandatory for the MVP: printed photo, photo on a phone screen, pre-recorded video played on a screen.
  • Level 2, desirable: paper masks with cutouts, deepfakes played on a screen.
  • Level 3, explicitly out of the MVP: injected camera streams and 3D silicone masks.

The rule attached to it: every design decision must be justified against this list. If a signal doesn’t catch one of these attacks, it doesn’t go in.

Then came the golden rule that shaped everything after it:

The client never knows the script. The server reveals one step at a time.

sequenceDiagram
    participant U as 📱 Client
    participant S as 🛡️ Server
    U->>S: Open session
    Note over S: Draws a secret, random script
    U->>S: Live camera frames (WebSocket)
    S-->>U: Challenge 1 — and only challenge 1
    U->>S: Frames reacting…
    S-->>U: Challenge 2 — revealed just now
    U->>S: Frames reacting…
    Note over U: Never knows how many steps,<br/>what comes next, or the thresholds
    S-->>U: LIVE · SPOOF · INCONCLUSIVE

A pre-recorded video can’t answer a question that didn’t exist when it was recorded.

And, just as important, the proposal included a list of what not to build: no face detection from scratch, no face matching, no native app, no billing, no back office. Plus a warning I didn’t want to hear: build the attack bench before the product, not after. That’s the part most people skip, and it’s the part that decides whether you have a product or a demo.


Step 2: Research only concrete decisions

My second question was simple: what stack should the backend use, and what’s best for processing video?

I don’t do abstract research at this stage. I bring specific forks in the road and want opinionated answers. Some of the answers stung, because they went against the stack I use every day:

  • Kafka was the wrong tool. Video frames are ephemeral, high-frequency and worthless after 40 ms. Persisting them to disk just to read and discard them is latency and cost for nothing. A fire-and-forget bus (NATS core, no persistence) was the right fit.
  • A durable workflow engine was the wrong tool. It’s built for workflows that last minutes to days, not a 30 ms decision loop. A liveness session is disposable by design: if it fails, you start over.
  • WebRTC and H.264 would erase the signal. Inter-frame compression smooths exactly the subtle color differences the light challenge needs to measure. Plain JPEG per frame, no temporal dependency.

Then I made the call that became the backbone of the system:

Python only for real-time image analysis. Everything else in Go.

Go owns the WebSocket, the session state machine, the challenge script, the score fusion, the verdict and persistence. Python only measures. That sentence would later become a boundary enforced in code, but I’m getting ahead of myself.


Step 3: Not one mega-prompt, a sequence of gated prompts

My next message: “give me the prompts, in steps, for Claude Code.”

For simpler projects I used to write one big mega-prompt. For this one, the build became a sequence of prompts, each one a phase, with three rules:

  1. One prompt per agent session.
  2. Commit between every prompt.
  3. Don’t advance if the acceptance criterion doesn’t pass.
flowchart LR
    P0["Prompt 0<br/>scaffolding + CLAUDE.md<br/>NO logic"] --> G0{"✅ make up<br/>make test"}
    G0 --> P1["Prompt 1<br/>message contracts"]
    P1 --> G1{"✅ gate"}
    G1 --> P2["Prompt 2…N<br/>state machine, analyzers,<br/>fusion, hardening"]
    P2 --> GN{"✅ gate"}
    G0 -. "fails" .-> P0
    G1 -. "fails" .-> P1
    GN -. "fails" .-> P2

Prompt 0 was the one that paid off the most. It asked for no logic at all: just the repo structure, the infrastructure in Docker, and a CLAUDE.md documenting the architecture, the exact Go/Python boundary, the threat model, and the golden rule about the script. Its acceptance criterion was almost boring: make up starts the infrastructure and make test runs, even with zero tests.

That boring file is what stops an agent from reinventing your protocol in prompt 7.

Prompt 1 was the contracts, the only thing Go and Python share, defined before either side existed. And the anti-injection prompt had my favorite acceptance criterion of the whole project: clear documentation of the threat model covered and NOT covered.


Step 4: Delegate, gate by gate

Then I ran the prompts. The agent wrote the code; I acted as architect and reviewer.

What came out the other side

Here’s the platform the gated prompts produced. Two Go services own everything that decides; one Python service owns everything that looks at pixels:

flowchart TB
    subgraph CLIENT["📱 Client (browser)"]
        CAM["Camera<br/>480p · native aspect ratio"] --> ENC["JPEG encoder<br/>in a Web Worker"]
        PAINT["Challenge painter<br/>logs exactly when it painted"]
    end

    subgraph DECIDE["⚖️ Go — decides"]
        GW["Gateway<br/>WebSocket · auth · backpressure"]
        ORC["Orchestrator<br/>state machine · secret script<br/>score fusion · verdict"]
    end

    subgraph MEASURE["🔬 Python — only measures"]
        AN["Analyzer<br/>face landmarks · pose · gaze<br/>light response · pulse · texture"]
    end

    subgraph DATA["🗄️ State and evidence"]
        R[("Redis<br/>hot session state<br/>always with TTL")]
        PG[("Postgres<br/>results + append-only audit")]
        S3[("Object storage<br/>encrypted evidence<br/>deleted after 24 h")]
    end

    ENC -- "frames over WSS" --> GW
    GW -- "one challenge at a time" --> PAINT
    GW <--> ORC
    GW -- "fire-and-forget bus" --> AN
    AN -- "named numbers only" --> GW
    ORC --> R
    ORC --> PG
    ORC --> S3

Every session walks through the same state machine, and there’s no way back once it’s resolved. No resuming, no retrying with the same script:

stateDiagram-v2
    [*] --> Created: one-time token
    Created --> Calibrating: client connects
    Calibrating --> Challenge: baseline captured
    Challenge --> Challenge: next step revealed
    Challenge --> Evaluating: script finished
    Evaluating --> LIVE
    Evaluating --> SPOOF
    Evaluating --> RETRY: quality, not attack
    Created --> Expired: timeout
    Challenge --> Expired: timeout
    LIVE --> [*]
    SPOOF --> [*]
    RETRY --> [*]
    Expired --> [*]

And this is the life of a single frame, from the camera to a number:

sequenceDiagram
    participant C as 📷 Camera
    participant W as 🧵 Web Worker
    participant G as 🛡️ Gateway (Go)
    participant A as 🔬 Analyzer (Python)
    participant O as ⚖️ Orchestrator (Go)
    C->>W: New frame ready (not a timer!)
    W->>G: JPEG + capture timestamp
    Note over G: Bandwidth cap, not frame-rate cap.<br/>A late frame lies about when it happened.
    G->>A: Frame on the bus (lost = never retried)
    A->>A: Landmarks, pose, gaze, light response…
    A-->>G: Named measurements, no verdicts
    G->>O: Measurements for this window
    Note over O: Only O knows the script,<br/>the weights and the thresholds

The first client was a bare JavaScript test harness, no framework. Its real job wasn’t the UI. It recorded every session frame by frame, with timestamps, camera settings and events, so any measurement could be replayed offline with the exact production code. That harness turned out to be the most important “feature” of the MVP.

When the skeleton was running, the questions got harder. How do I catch masks? What does it take to pass the certifications the big providers have (iBeta Level 1 and 2, under ISO/IEC 30107-3)? One lesson from that research changed how I read every metric afterwards: rejecting too many real people disqualifies you just as much as letting attackers through. Security and usability are graded together.


Step 5: Measure, and let the data humble you

It was past midnight, and the numbers were perfect.

Every time my screen flashed, my face answered. Twelve measurement windows, twelve near-perfect correlations. I remember leaning back and thinking: it works.

The next day I ran the exact same session under different real-world conditions.

Six windows. Six failures.

My first instinct was to open the agent and type “fix it.” That’s exactly the instinct this recipe exists to stop. An agent told to “fix it” will find a way to make the numbers green, and in security, green numbers lie.

Three sources of truth

So I stopped trusting any single kind of test. The platform is tested with three sources of material, from cheapest to most expensive, and each one answers a different question:

flowchart LR
    F["🤖 Synthetic mannequin<br/><i>cheap · every code change</i><br/><b>Does the pipeline run<br/>end to end?</b>"]
    R["🎥 Real recordings<br/><i>medium · when a real session fails</i><br/><b>Does it measure well with<br/>a real face and camera?</b>"]
    D["🎯 Attack datasets<br/><i>expensive · when a detector or threshold changes</i><br/><b>Does it separate a person<br/>from a photo, screen or mask?</b>"]
    F --> R --> D

🤖 The mannequin. A head drawn in 3D, in several skin tones, generated by the same scene code that validates the analyzer’s signals. It responds to calibration, to light, to turning sideways and to moving closer. It plays full sessions against the real gateway, with the exact modules the browser loads. Only two things are swapped: the camera, for synthetic frames, and the screen, for a painter that logs when it painted.

It has limits, and I love what they reveal. The mannequin can’t raise its chin or look at a dot, so the bench simply keeps playing sessions until it draws a script the mannequin can answer. Why not just fix the script for testing? Because the seed never leaves the server. Not even for my own test bench. The golden rule applies to me too.

The synthetic frames the bench plays as a camera, one per challenge colour

A synthetic session that passes proves the pipeline works. It proves nothing about real people. For that, there’s the second source.

🎥 Real recordings. This is what I used the most, in a five-step cycle:

flowchart LR
    S1["1️⃣ Tunnel up<br/>phone reaches client<br/>and gateway over HTTPS"] --> S2["2️⃣ Play a session<br/>on the phone,<br/>recorded frame by frame"]
    S2 --> S3["3️⃣ Read it live<br/>verdict + detector table<br/>+ gateway log"]
    S3 --> S4["4️⃣ Recompute offline<br/>any measurement, with<br/>production code"]
    S4 --> S5["5️⃣ Check controls<br/>real vs reversed<br/>vs no-challenge"]
    S5 -. "hypothesis survives" .-> FIX["🔧 Change code<br/>+ test with real values"]
    FIX -.-> S2

Each recording keeps everything: which challenges were revealed and when, the capture timestamp of every frame, exactly when each color was painted, camera settings, dropped frames and costs, followed by the JPEGs exactly as they were sent. That means I never have to ask myself to repeat a session to test an idea. At one point, the photometry of 13 recordings was extracted once and then used to test estimator after estimator without detecting a single face again.

Step 5 is the heart of the method. Every measurement is computed three times:

flowchart TB
    M["📏 One measurement"] --> A["✅ Real sequence<br/>what we want to measure"]
    M --> B["🔄 Reversed sequence<br/>if it scores the same, the algorithm<br/>fits any wave and measures nothing"]
    M --> C["🤫 Still segment, no challenge<br/>if it scores here,<br/>we're measuring noise"]
    A & B & C --> V{"Real beats<br/>both controls?"}
    V -- "Yes" --> OK["Touch the code"]
    V -- "No" --> NO["Into the graveyard"]
The colour response of a real session against the moment each colour was painted, with its two controls

🎯 Attack datasets. Public datasets of real people and attacks: masks, replays, bent and cut-out photos, screens. They only test the passive detectors (texture, moiré, banding), because a photo in a dataset can’t respond to a light challenge or look at a dot. Results are always reported per attack type, never aggregated, and thresholds are chosen on one half and measured on the other, split by subject. The experiment is repeated with several seeds, because how much the result moves between seeds is part of the result. Every run is written down: what was measured, and how far the conclusion reaches.

The loop

Put together, every change goes through the same loop:

flowchart TD
    M["📏 Measure a real session"] --> H["💭 Form a hypothesis"]
    H --> R{"Does it beat<br/>the control?"}
    R -- "No" --> L["📕 Log it as REFUTED<br/>nobody tests it again"]
    R -- "Yes" --> CH["🔧 Change code or profile"]
    CH --> T["🧷 Add a test with the real case"]
    T --> M

The graveyard

The spec has a section most people would be embarrassed to keep: everything that didn’t work, and why. It’s the most useful page in the repo. Here’s a slice of it:

🪦 What I tried💥 What actually happened
Capture frames on a timerHalf the frames were copies of the previous one: new timestamp, old pixels.
Compress frames as WebP28 ms per frame versus 1 ms for JPEG. Frame rate dropped from 21 to 7 fps.
Capture at 720pIt needed 20.7 Mbit/s of upload. One network hiccup swallowed an entire challenge.
Rate-limit frames by arrival timeNetwork jitter bunched frames together, and the gateway threw away 72 of 172 legitimate ones.
Optimize the face filter, move the codec to a Worker, poll twice as fastNone of them was the bottleneck. Total gain: less than 1 fps. The real limit was the camera and the link.
Stretch every frame to 640×480Phones deliver in portrait. Every face came out 1.78× wider, breaking head pose, eye openness and gaze, all at once.
Longer light challengesCorrelation collapsed from 0.97 to 0.05. The camera’s auto-adjustment quietly erased the very effect I was measuring.
A model 7× bigger than the one I keptIt didn’t perform better on my bench, despite a better published score.
A fine-tuned vision transformerWorse than the small model, and worse than its own un-tuned version.
Average two anti-spoofing modelsThey looked at the same crop, so the system counted the same evidence twice. Worse than one alone.
Screen glare as a detectorIt let 61% of attacks through.
Reuse a threshold from one dataset on anotherIt rejected two out of three real people.
Split the test set by fileThe same face landed on both sides. I was measuring memory, not generalization.
Pulse measured in absolute valuesWhite-balance drift sank its accuracy (AUC) from 0.71 to 0.42, worse than a coin flip.
Change the user’s environment to get a cleaner signalI rejected it. If a product only works when users adapt to it, it isn’t a product.
Frames per second captured in 34 recorded sessions against the 30 the server asked for

Now read that table again and count how many of those failures hurt real users, not attackers. That’s the part nobody warns you about. I’ll come back to it.

And the midnight failure? Several alternative estimators were tested against controls. None of them recovered the signal. So instead of letting an agent overfit its way to green, it’s written in the spec: known limit, the evidence behind it, and the mitigation plan.

The same goes for what hasn’t been proven yet. The spec says it out loud: real physical attacks in front of the camera, a printed photo, a replay on a screen, a mask, are still untested live. Until they are, this is a system being refined, not a system ready for certification.

In security, documenting your limits isn’t weakness. Hiding them is.


What this project added to my recipe

Looking back, the recipe I started with had four holes. This project filled them, and they’re now part of how I build everything.

1. Done means the attack fails

“Done” can’t be a feeling. On this project it became four sources of proof:

ProofWhat it covers
🧪 423 automated tests245 in Go, 159 in Python, 19 in the JS client
🎯 Attack benchTwo public datasets, including 8,126 images of 30 subjects, split by subject
📱 59 real phone sessionsRecorded, then recomputed offline with production code
🕳️ What’s still untestedWritten down, not hidden

The number isn’t the point. What the tests guard is. The most important ones read like the golden rules of the threat model, because that’s exactly what they are:

  • TestNoMessageLeaksTheScript
  • TestPublicAPINeverExposesFutureSteps
  • TestScriptStaysUnpredictable
  • TestDecisionTable, which proves that no quality path can ever end in a rejection

Many of them are property tests that run over thousands of random seeds. That’s how one “clever” design died: drawing a random value and retrying until it was valid looked fine in a demo, and failed across 2,000 seeds, because sometimes no valid value existed at all. No human reviewer would have caught that. A property test did, in seconds.

Plus the rule that caught more false progress than any code review: every measurement gets a control. Run it on a reversed challenge sequence, or on a stretch where nothing happened. A measurement that scores the same on the control measures nothing.

2. Write down decisions and what you ruled out

Chat history is where good decisions go to die. Now every decision lives in the spec next to the measurement that justifies it, and refuted hypotheses are written down too, so neither I nor the agent wastes a week re-testing them.

3. Phase 0 is privacy and compliance

Biometric data is special-category data in many jurisdictions, including where I live. It would have been easy to leave that for the end. Instead, it’s phase zero:

flowchart LR
    CAP["📷 Captured"] --> HOT["⚡ Redis<br/>session state<br/>expires on its own"]
    CAP --> EV["🗝️ Evidence<br/>encrypted on the client"]
    EV --> DEL["🗑️ Deleted<br/>after 24 h"]
    CAP --> AUD["📜 Audit trail<br/>append-only<br/>no updates, no deletes"]
    CAP --> REC["🧬 Dev recordings<br/>one machine · never versioned"]
    REC --> DEL2["🗑️ Deleted when<br/>the tuning is done"]
  • ⏳ Session state always has a TTL.
  • 🔒 The audit trail is append-only, enforced by database triggers, not by good intentions.
  • 🗝️ Evidence is encrypted client-side and deleted after 24 hours.
  • 🧬 Session recordings never get versioned and never leave the machine.
  • 🕵️ A pentest mode shows the full breakdown of every verdict, for testing only. In production, a rejected user sees exactly two words: “try again.” Anything more hands the attacker an oracle.

4. Say what you won’t build, and enforce boundaries in code

The spec says plainly what’s out of scope: injected camera streams and real-time deepfakes need a native SDK with device attestation, and this web MVP doesn’t claim to stop them.

And “Go decides, Python measures” stopped being a sentence and became a wall:

flowchart LR
    C["📱 Client"] -- "live frames" --> G["🛡️ Gateway · Go"]
    G <--> O["⚖️ Orchestrator · Go<br/>state machine · verdict"]
    G -- "frames" --> A["🔬 Analyzer · Python<br/>vision + signal"]
    A --> K{{"🧱 Codec"}}
    K -- "measurement = 0.87 ✅" --> G
    K -. "is_live = true ❌ REJECTED" .-> X["🚫"]

The codec between them rejects any signal named like a verdict: is_live, spoof_score, anything that smells like judgment. An agent working on the Python side literally cannot make it decide, no matter how helpful it’s trying to be.

A rule an agent can’t break is worth ten rules it’s asked to remember.


The rules that protect people

One warning from the very first proposal stuck with me: if you don’t calibrate each person against their own baseline, your rejection rate will be brutally worse for darker skin. That’s a product defect, not a detail.

Remember the graveyard? Here’s the part of it I didn’t show you yet. Some of the most painful failures had nothing to do with attackers:

  • 👁️ A fixed “eyes open” threshold meant a user whose eyes are naturally more closed lost every single frame.
  • 🎯 Measuring gaze against a “rest position” failed in three different ways, and one of them flagged a legitimate user as fraud.
  • 🚨 Letting the light-challenge correlation veto on its own accused real sessions of fraud.
  • 🪪 A hard floor on identity continuity rejected two real people. Three redesigns of that statistic failed too.

None of these would show up in a demo with the developer’s own face. They show up when the system meets people who aren’t you.

In KYC, a false rejection isn’t a statistic. It’s a person locked out of a bank account. So three principles carry the same authority as the security invariants:

  1. 🤷 What couldn’t be measured doesn’t vote. Missing isn’t guilty.
  2. 🌫️ Bad quality is not an attack. Blur or bad lighting can end in a retry, never in a rejection.
  3. 💓 The pulse can acquit, but never accuse. Remote pulse detection depends on skin tone, so it can add confidence but never count against anyone.

No agent optimizing for a better score would ever invent these. That’s exactly why they have to be written down.


The recipe, in one page

  1. Ideate with no code and no demos. Start with the attacker (or the user), the golden rules, and what you won’t build.
  2. Research only concrete forks. Ask for opinions, and be ready for answers that go against your usual stack.
  3. Turn it into gated prompts. Prompt 0 is scaffolding plus CLAUDE.md, no logic. One prompt per session, commit between, never advance past a failing gate.
  4. Delegate. Be the architect and the reviewer.
  5. Measure and refute. Controls for everything, and every lesson goes back into the rules.

And the spec skeleton every new project of mine starts with:

text
# CLAUDE.md — <project>

## Purpose (one sentence)
## Threat model / golden rules            ← before any code
## Architecture and ownership             ← who decides, who only measures
## Decisions log                          ← decision + the measurement behind it
## Graveyard: what failed and why         ← so nobody tests it again
## Phase 0: non-functional                ← privacy, retention, audit, secrets, observability
## Phases                                 ← each: build X, done when <measurable check>
## Out of scope                           ← said out loud
## Boundaries enforced in code            ← and where

Back to that first message

“NO CODE. NO DEMOS.”

It looks like impatience. It’s actually the whole method. When agents write the code, your rules become the real attack surface, and every ambiguity is a place where the agent will confidently choose “works” over “safe.”

So write the rules first. Write them like a security engineer. Measure them against controls. Make the important ones impossible to break.

I didn’t write the code.

But I’d sign every rule it follows.

🔗 Read the code, the spec and the graveyard on GitHub: github.com/edisonpaul4/face-liveness-platform.