Ainoverse

Ainoverse

Reliability, failure handling and acceptance criteria

Before you call an automation finished, you need cases with known answers. The sample cases below are illustrative inputs, not measured results.

Acceptance criteria should be written before the first run. For the draft-only pilot, a natural set is: every factual sentence in the draft traces to a line in the CV; every requirement in the posting is either matched or explicitly listed as a gap; zero outbound messages occur without a recorded human approval; a failing tool halts the run after at most two total attempts (one retry) and names the step.

Sample case one: the CV and posting align on four of five requirements. Expected outcome — four overlaps appear in the match step, the fifth appears as a labelled gap, the draft claims only the four, and the run proceeds to review. Sample case two: the drafting tool returns a malformed response. Expected outcome — one retry, a second malformed response, the run halts, and the reviewer receives a note identifying the drafting step as the failure point; no partial draft is presented for approval.

Illustrative test cases and expected outcomesMissing credential → show gapInvented CV fact → reject draftDuplicate draft → review, nosendFailed tool → bounded retriesApproval missing → do not send
Illustrative test cases and expected outcomes

Sample case three: the job posting was taken down between selection and drafting. Expected outcome — the fetch fails validation, the run stops before producing a draft, and the reviewer is told the listing could not be retrieved. Sample case four: the reviewer edits a draft line to add a detail not present in the CV. Expected outcome — the edited line is flagged in review as untraceable, since the faithfulness check compares the final text, not only the generated text.

None of these outcomes is a success percentage. They are expected behaviours under specified inputs, which is what acceptance criteria can honestly provide. Reliability comes from the failure paths being as explicit as the happy path — and from the approval gate remaining the last thing before anything external happens.

  • Case A — 4 of 5 requirements matched: four overlaps, one labelled gap, draft claims only the matched four.
  • Case B — tool malformed twice: halt after two attempts, no partial draft presented, failure step named.
  • Case C — listing unavailable: fetch fails validation, run stops before drafting, reviewer notified.
  • Case D — reviewer adds an untraceable line: final-text check flags it before approval is recorded.

Further reading