Romeo Apps owned proof / build to proof

Synthetic benchmark / rubric → receipt

Good evaluation makes the boundary visible.

A small, inspectable lab for turning AI engineering outputs into task-specific rubrics, explicit terminal states and the next safe action.

RomeoApps-owned synthetic proof. No credentials, network requests, customer records or external writes.

03synthetic tasks
15rubric points
00network requests
01matrix receipt

01 / replay a task

Grade the output. Keep the next action.

Each case shows the prompt, candidate output, rubric and terminal state. The matrix is deliberately small so every judgment can be inspected rather than hidden behind a single aggregate score.

browser reliabilitysynthetic

Choose a locator that survives UI change

candidate output
rubric / max 5— / 5
    event tracelocal replay
    1. Waiting for a local replay.

    02 / acceptance matrix

    A score without a terminal state is not enough.

    The cases cover a robust implementation, an underspecified contract and an unsupported release claim. Their expected outcomes are different on purpose.

    FixtureJudgment seamExpected stateWhy
    Browser locatorsemantic selectorPASSRobust target, bounded action.
    API contractmissing behaviorHOLDDo not invent production semantics.
    Release notemissing evidenceREVISEAssertion is not a receipt.

    03 / method

    Task → rubric → terminal state → receipt.

    This is the smallest useful evaluation loop: name the task, make quality observable, stop safely when a requirement is missing, and leave the next question explicit.

    01Taskrealistic seam, bounded input
    02Rubricnamed criteria, points
    03DecisionPASS / HOLD / REVISE
    04Receiptevidence + next action

    04 / truth boundary

    Useful proof stops before private access.

    Verified by this page

    • Three deterministic synthetic evaluation fixtures.
    • Rubric rows with explicit point values and expected states.
    • Local matrix replay with a stable receipt id.
    • Zero network requests and zero external writes.

    Not claimed

    • Client, partner-lab or production model results.
    • Benchmark-suite accreditation or model-provider affiliation.
    • Selection, paid work, revenue or customer ROI.
    • That a synthetic task substitutes for buyer-owned acceptance.

    Fit before payment

    Bring one evaluation seam that needs a reliable answer.

    For a real engagement, start with a written task, safe evidence, a named acceptance owner, the payment route and the boundary around private data. I can turn that into a bounded benchmark or reliability slice.

    Request a written fit