Choose a locator that survives UI change
Romeo Apps
Synthetic benchmark / rubric → receipt
A small, inspectable lab for turning AI engineering outputs into task-specific rubrics, explicit terminal states and the next safe action.
RomeoApps-owned synthetic proof. No credentials, network requests, customer records or external writes.
01 / replay a task
Each case shows the prompt, candidate output, rubric and terminal state. The matrix is deliberately small so every judgment can be inspected rather than hidden behind a single aggregate score.
02 / acceptance matrix
The cases cover a robust implementation, an underspecified contract and an unsupported release claim. Their expected outcomes are different on purpose.
| Fixture | Judgment seam | Expected state | Why |
|---|---|---|---|
| Browser locator | semantic selector | PASS | Robust target, bounded action. |
| API contract | missing behavior | HOLD | Do not invent production semantics. |
| Release note | missing evidence | REVISE | Assertion is not a receipt. |
03 / method
This is the smallest useful evaluation loop: name the task, make quality observable, stop safely when a requirement is missing, and leave the next question explicit.
04 / truth boundary
Fit before payment
For a real engagement, start with a written task, safe evidence, a named acceptance owner, the payment route and the boundary around private data. I can turn that into a bounded benchmark or reliability slice.
Request a written fit