Skip to main content
Who this is for: recruiters, hiring managers, and integrity teams reading the output of HeyMilo’s interview cheat detection (V2).

Read a result in 30 seconds

Every analysis returns one overall verdict plus a list of individual findings (called breaches). Start at the top:
Golden rule: the system assists a human decision — it does not auto-fail candidates. Treat Review and any single finding as “a human should look at this clip,” not “this candidate cheated.”

Reading the risk_summary

This is the top-of-report banner — the verdict and its calibration.
This is a recommendation, not an automatic decision. The final call is always yours.

How the overall call is computed

The verdict is based on how much of the interview was compromised, not just the count of findings:
Why duration, not count? A single 2-second false blip shouldn’t fail a candidate, and ten tiny flags shouldn’t outweigh one sustained, serious breach. Measuring the share of the session keeps the verdict proportional. You control the reject cutoff via rejection_threshold.

Calibration fields

Important: compromised_percentage counts only high-severity findings. Medium and low findings are real but advisory — they appear in the breakdowns and breaches[], but don’t inflate this headline number. So “0% compromised” can still come with medium/low findings worth a glance.

Composition fields

  • severity_breakdown — how much time fell into high/medium/low buckets (count + duration + % each). Always shows all three buckets, even at zero.
  • category_breakdown — which kinds of issues appeared (identity, external assistance, etc.), biggest offender first.

Reading an individual breach

Each finding looks like this:

The severity ladder

Most “uncertain” findings land at medium on purpose. The system surfaces plausible signals at medium so a human can decide, rather than hiding them or over-calling them high.

About confidence

confidence (0–1) is a display signal of how strong the underlying detection was — it does not drive the verdict (severity and duration do). Read it as a rough “how strong is the signal,” not a probability of cheating.

Breach catalog — what each finding means

Findings are grouped into five integrity dimensions (categories). Below is what each finding means, what triggers it, and how to interpret it.

1. Identity Authenticity Integrity

Is the person on camera really the enrolled candidate? identity_authenticity_risk — the face on camera doesn’t match the enrolled candidate, or the identity changes mid-interview.
  • Triggers on: a face that clusters differently from the enrollment photo, or an abrupt person-switch.
  • Interpret: high = a clear mismatch or swap; medium = ambiguous (pose, lighting, low video quality). Genuine network/camera-quality artifacts are meant to be discarded — if you see this at medium, check whether it’s just a bad frame.

2. Presence and Participation Integrity

Is the candidate continuously present and observable? candidate_absence_or_unverifiable_presence — the candidate leaves frame or becomes unidentifiable for a sustained stretch.
  • Triggers on: the face being missing/untrackable for longer than the configured floor.
  • Interpret: high = prolonged total absence; medium = partial or hard-to-verify presence. A brief look-away is meant to be ignored.

3. External Assistance Integrity

Is the candidate getting outside help? unauthorized_aid_object_present — a phone, tablet, second screen, book, or notes is visible/usable.
  • Triggers on: detection of a monitored object class (configurable).
  • Interpret: presence alone is medium at most; high requires visible use (reading from, typing on, or repeatedly glancing at it).
external_voice_assistance_detected — a non-candidate voice is coaching or supplying answers.
  • Triggers on: a second speaker overlapping the candidate’s answer windows.
  • Interpret: high = the voice clearly shapes the answer; medium = another voice is present but its role is unclear (could be background noise).
offscreen_reference_pattern — sustained gaze off-camera while answering (a second monitor, side notes, prompts).
  • Triggers on: gaze held beyond an angle/duration threshold — distinct from natural thinking glances.
  • Interpret: high = appears to be reading from an off-camera source (often corroborated by other findings); low/medium = sustained but ambiguous. Brief or natural glances are discarded.
multiple_people_presence — more than one person in frame.
  • Triggers on: a second person detected with persistence.
  • Interpret: high = a second person appears to assist; medium = present but role unclear. Reflections, posters, and people passing by are meant to be discarded.
covert_response_sync_pattern — a repeated long-pause-then-immediate-answer rhythm suggesting an external prompt loop.
  • Triggers on: the delay between the question ending and the answer starting exceeding the configured limit.
  • Interpret: high = timing clearly synced to an external cue; medium/low = suspicious but explainable by normal thinking time. Slow-but-genuine thinkers can trip this — tune the delay for your role.

4. Session Trust Integrity

Is the session itself trustworthy (not spoofed or looped)? session_tampering_suspected — video feed spoofing, looping, or audio/video desync.
  • Status: On the roadmap — not yet active in detection. The category exists for forward-compatibility; you won’t see findings for it today.

5. Response Authenticity Integrity

Are the answers the candidate’s own, in real time? synthetic_response_generation_suspected — an answer reads as AI-generated.
  • Triggers on: the AI-text probability of a qualifying answer crossing the threshold.
  • Interpret: high = strongly AI-like; medium = polished but could be genuine. The finding cites the answer in evidence.metadata.
paraphrased_response_suspected — an answer reads as AI-reworded (the candidate’s draft polished by an AI assistant).
  • Triggers on: the paraphrase probability crossing its (separate) threshold.
  • Interpret: distinct from synthetic — “AI helped rewrite” vs “AI wrote it.” Tune the two thresholds independently.

Policy parameters — tuning sensitivity

These knobs let you tune the system to your role and risk tolerance. They live in your policy (set per role/job). Anything you don’t set uses the default below.

Global knobs (apply everywhere)

These can also be set per category or per breach type — a more specific setting overrides a broader one.

Per-breach knobs

You can also turn whole categories or individual breach types on or off via their enabled flag — e.g. a voice-only screening can disable all video breaches.

Worked examples

Example A — “Proceed with a note”

Read it as: No high-severity time, so the overall call is Proceed. There’s one low-severity gaze finding — the candidate glanced away, but it’s weak and unconfirmed. Glance at the clip if you like; not a basis to act.

Example B — “Review”

Read it as: No high-severity time (so 0% compromised), but two medium findings pushed the combined share past the review line. A phone was visible (use unconfirmed) and answers came after long pauses. Watch both clips — together they may tell a story, or both may be innocent.

Example C — “Reject”

Read it as: High-severity findings cover 18% of the session — over the reject threshold. A second voice was supplying answers and the face didn’t match enrollment. Confirm the clips, then disqualify if they hold up.

FAQ

Q: Why was something flagged but the verdict is still “Proceed”? Because the verdict measures the share of the session covered by high-severity findings against your rejection_threshold. A short or medium/low finding won’t move the overall call — but it’s still listed in breaches[] for your awareness. Q: Why is compromised_percentage 0% when there are findings? That number counts high-severity only. Medium/low findings are intentionally excluded so the headline reflects confirmed serious compromise. Check severity_breakdown and breaches[] for the rest. Q: What does “medium” really mean — did they cheat? Medium means plausible but unconfirmed. The reviewer saw something that could enable cheating but couldn’t be sure from the evidence. It’s an invitation to look, not a conclusion. Q: Can the system reject a candidate on its own? No. It produces a recommendation. A human always makes the final call.

What this system does not do

  • Does not make hiring decisions — it surfaces evidence for a human
  • Does not prove intent — a flag is “worth reviewing,” not “guilty”
  • Does not catch everything — coverage is bounded by what’s visible in sampled evidence
  • Is not a biometric identity service — identity findings compare against the enrollment you provided, for integrity purposes only
Use it as a force-multiplier for reviewers: it points you to the 30 seconds worth watching in a 45-minute interview, explains why, and tells you how confident it is — so your time goes to judgement, not scrubbing.

Back to setup