Skip the home page. Ask them to log in, pick one genuinely messy incident from last month, and walk it through in order. The order is the test. Almost nothing survives it, and you will know inside ten minutes.
Start where work arrives. Not the chat box — the place a request lands when nobody decided to make it a request. A forwarded mail, an attachment, a portal submission, a photograph from a site. Ask what happens to that arrival before a human types anything. If the answer is that somebody reads it and creates a row, the product has not observed anything. It has a form. Worth counting, while you are there: the requests that never became rows at all. That number is the demand a company cannot see, and no board has ever held it.
Then ask where it is written down. You want an append-only record of the arrival: what came in, from whom, when, through which channel, reconstructable later. Not a comment thread. Not a summary a model produced. The distinction to press on is evidence versus interpretation — the mail said three weeks, the system inferred a delay. Both belong in the record and they must not be the same row. Products that collapse them cannot ever explain a wrong conclusion, because by then the inference looks like a fact.
Now ask for it. One durable situation, with a timeline: observations, predictions, proposals, decisions, executions, verifications, all bound to the thing itself rather than to whoever happened to be in the thread. This is where most demos go quiet. Across our whole corpus there is no vendor marketing a case object as a first-class surface. You will be shown a ticket, a workflow run, or a chat thread with memory attached. A ticket closes. A run ends. A thread belongs to its participants. A situation outlives all three, and it is the single largest hole in this market.
Ask what the system predicted, and when. Not what is late — what breaks next. The honest version has a timestamp, because the whole value is arriving early. A prediction that shows up in the same week as the consequence is a status report wearing a hat.
Ask to see a proposal. A real one: intent, the evidence behind it, the class of effect it would have, and — for anything that commits money or time — the lead time and the schedule consequence, on the card, before anyone can approve. The strongest peer walkthrough in our set runs an invoice through validate, route, auto or flag, then posts it to the finance system. Clean, deterministic, human only at the decision node. Another peer shows the actual script next to an Approve and Run button, which is the clearest write gate anybody in this corpus has drawn. Ask for that, and watch whether approve is disabled when the consequence fields are missing.
The gate, and take your time here. Proposed, policy-checked, auto-allowed, awaiting a human, clarifying, approved, rejected, escalated, stalled, timed out, executing, verify-pending, verified or failed, cancelled. That is fourteen states, and the two that separate a product from a badge are stalled and verify-pending. Stalled is not slow. One peer shows an approval overdue with a named person unresponsive for thirty-six hours and a tile reading SLA at risk — the closest thing in the market, and still a threshold on a clock rather than a distinct state with its own treatment. Ask the demo to show you the queue sorted with stalled at the top. Watch what they click.
Ask what the system did. Execution should name the external system it touched, the same way a good pause names what it refuses to touch without permission: credentials, production, files, anything crossing the company boundary. You approve at the boundary, the machine runs inside it.
Then ask the question that ends most demos. How does the product know the intended effect actually happened? Not that the call returned successfully — that the purchase order exists, that the date moved, that the system of record on the other end agrees. This is the bricked-up door. Verify is missing almost everywhere, including in the best peer walkthroughs, and a loop that ends at a successful call is not closed, it is abandoned politely.
Finish in the audit. Pick the decision from four steps ago and reconstruct it: what was proposed, which policy fired, who approved, on what evidence, using which version of the routine. If editing a routine silently rewrote the history of runs that used the old one, stop. That single defect makes every other claim on the page unprovable.
What is still missing, everywhere, is worth saying plainly. The situation object. A front page that is exposures and waiting decisions rather than an activity feed. Stalled held apart from slow. A verify surface. A visible emergency stop that revokes rather than pauses.
Nine steps. One incident. Bring your own — the one they pick for you will be the one that works.
Talk soon — bring one messy incident.

