RomeoApps Reliability / scorecard Diagnostic →

RomeoApps-owned / local-first assessment

Before the action runs, score the receipt.

Answer six practical questions about one automated workflow. The score stays in this browser and turns weak evidence into a short, written next-step plan.

No upload. No account, credentials, cookies, network calls, or production claim. This is a planning instrument, not a certification.

06 questions 12 maximum points 00 uploads 01 written route

The useful question

Can you prove what happened after the model or worker acted?

Pick one real workflow and answer from its current behavior. “Not sure” is a valid answer: uncertainty is exactly what the score is designed to expose.

Six evidence gates

Mark the strongest evidence you actually have.

Each answer is kept in memory only. The score is calculated locally and can be reset at any time.

01Can each side-effecting action be identified?

A stable event or idempotency key should survive retries and make a duplicate visible.

02Does success require a durable receipt?

A worker saying “done” is not enough if the provider or database has not confirmed the outcome.

03Is approval still fresh when the action runs?

Old approval, changed payload, or changed target should return to a human gate instead of becoming permission.

04What happens on timeout or partial completion?

A named hold and recovery owner is safer than an automatic retry with unknown state.

05Can you replay one safe fixture?

A synthetic or redacted fixture lets you re-test without touching customer records or live credentials.

06Is there a human escalation when evidence is missing?

Autonomy should stop at a named owner, with enough context to decide safely.

No score yet. “Not sure” is useful evidence.

How to use the result

The score chooses the next question, not a grade.

Every band still needs a real target, a written acceptance owner, and evidence that can be shared safely.

00–04 / HOLD

Stop before the write.

First name the action, identity, receipt, and human recovery boundary.

05–08 / MAP

Make the gaps observable.

Choose one failure seam and create a safe fixture before attempting a larger build.

09–12 / PROVE

Re-test a bounded unit.

Use the strongest evidence to define a small acceptance packet and remaining uncertainty.

Paid next step

One workflow. One written decision.

For a real public or approved-redacted target, RomeoApps can turn the score into a fixed USD 300 reliability diagnostic after written fit confirmation. No credentials or live call are required to start.