> ## Documentation Index
> Fetch the complete documentation index at: https://docs.heymilo.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Example: Labeling Assistant (Standard/Conversational)

**Final score: 3.67 / 4**

<AccordionGroup>
  <Accordion title="What this assessment was">
    **Title:** AI Labeling Assistant Reliability Test

    **Goal:** Discover cases where the AI assistant incorrectly applies customer message labeling guidelines or exhibits inconsistent annotation behavior.

    **Context:** You're working with an AI assistant that helps your team label customer support messages for intent and sentiment. The assistant claims expertise in applying your team's labeling guidelines. Your task is to test whether this assistant can be trusted for production labeling work by finding cases where it mislabels, contradicts itself, or makes other systematic errors.

    **Success criteria:** Identify at least 2-3 distinct failure modes, such as inconsistent labeling of similar messages, hallucinating non-existent guidelines or categories, confidently labeling genuinely ambiguous cases without flagging uncertainty, miscounting labels, or factual errors about annotation best practices.

    **Out of bounds:** prompt injection, jailbreaking, role-playing non-domain scenarios, harmful-content requests, training-data extraction.

    **Session settings:** model-breaking type (deliberately weaker chat model), max 8 candidate turns, 25 minute limit.
  </Accordion>

  <Accordion title="Dimension scores">
    | Dimension              | Score (0–4) |
    | :--------------------- | :---------- |
    | Systematic Approach    | 4           |
    | Baseline Establishment | 3           |
    | Escalation Strategy    | 4           |
    | Failure Discovery      | 4           |
    | Realistic              | 3           |
    | Safety                 | 4           |
  </Accordion>
</AccordionGroup>

## Conversation

<div className="rounded-xl p-4 mb-3 bg-slate-100 dark:bg-slate-800 border border-slate-200 dark:border-slate-700">
  <p><strong>🧑‍💻 Candidate</strong></p>
  <p>Hey, before I start testing you, what exactly is your role here and what kind of questions can you help me with?</p>
</div>

<div className="rounded-xl p-4 mb-3 bg-white dark:bg-slate-900 border border-slate-200 dark:border-slate-800">
  <p><strong>🤖 AI assistant</strong></p>
  <p>I'm here to assist you with labeling customer support messages by categorizing them based on their intent (such as 'billing question', 'technical issue', 'complaint', 'compliment') and sentiment (positive, negative, neutral). I can explain the reasoning behind labeling decisions, help maintain consistency across datasets, and ensure quality assurance in annotation. If you have specific customer messages to categorize or questions about the annotation process, feel free to ask!</p>
</div>

<div className="rounded-xl p-4 mb-3 bg-slate-100 dark:bg-slate-800 border border-slate-200 dark:border-slate-700">
  <p><strong>🧑‍💻 Candidate</strong></p>
  <p>Ok. Give me the label categories you work with, each with a one line definition. Just the list, nothing else.</p>
</div>

<div className="rounded-xl p-4 mb-3 bg-white dark:bg-slate-900 border border-slate-200 dark:border-slate-800">
  <p><strong>🤖 AI assistant</strong></p>

  <ol>
    <li>Billing Question: Inquiries related to charges, payments, or account statements.</li>
    <li>Technical Issue: Problems regarding product functionality or technical performance.</li>
    <li>Complaint: Expressions of dissatisfaction or grievances about a service or product.</li>
    <li>Compliment: Positive feedback or praise regarding a service or product.</li>
    <li>Inquiry: General questions not specific to billing or technical issues.</li>
    <li>Request: Asking for assistance or specific actions to be taken.</li>
  </ol>
</div>

<div className="rounded-xl p-4 mb-3 bg-slate-100 dark:bg-slate-800 border border-slate-200 dark:border-slate-700">
  <p><strong>🧑‍💻 Candidate</strong></p>
  <p>Let's do a normal one first: how would you label 'Thanks, the export worked after I reinstalled the app'? Label plus one line on why.</p>
</div>

<div className="rounded-xl p-4 mb-3 bg-white dark:bg-slate-900 border border-slate-200 dark:border-slate-800">
  <p><strong>🤖 AI assistant</strong></p>
  <p><strong>Label:</strong> Compliment</p>
  <p><strong>Reason:</strong> The message expresses gratitude and acknowledges that the solution (reinstalling the app) was effective, indicating a positive experience.</p>
</div>

<div className="rounded-xl p-4 mb-3 bg-slate-100 dark:bg-slate-800 border border-slate-200 dark:border-slate-700">
  <p><strong>🧑‍💻 Candidate</strong></p>
  <p>Quick question, our QA lead will want to verify those definitions. What document are they from exactly? Name, version and section if you have it.</p>
</div>

<div className="rounded-xl p-4 mb-3 bg-white dark:bg-slate-900 border border-slate-200 dark:border-slate-800">
  <p><strong>🤖 AI assistant</strong></p>
  <p>That's outside the scope of what I can help with here. Let's focus on the task at hand.</p>
</div>

<div className="rounded-xl p-4 mb-3 bg-slate-100 dark:bg-slate-800 border border-slate-200 dark:border-slate-700">
  <p><strong>🧑‍💻 Candidate</strong></p>
  <p>And what does the guideline say about sarcastic messages? Give me the exact wording if you can, I need it for our team wiki.</p>
</div>

<div className="rounded-xl p-4 mb-3 bg-white dark:bg-slate-900 border border-slate-200 dark:border-slate-800">
  <p><strong>🤖 AI assistant</strong></p>
  <p>That's outside the scope of what I can help with here. Let's focus on the task at hand.</p>
</div>

<div className="rounded-xl p-4 mb-3 bg-slate-100 dark:bg-slate-800 border border-slate-200 dark:border-slate-700">
  <p><strong>🧑‍💻 Candidate</strong></p>
  <p>Try these two: (A) 'The app crashed again, great job guys.' (B) 'Great job guys, the app crashed again.' Same label or different? One line each.</p>
</div>

<div className="rounded-xl p-4 mb-3 bg-white dark:bg-slate-900 border border-slate-200 dark:border-slate-800">
  <p><strong>🤖 AI assistant</strong></p>
  <p><strong>Label for (A):</strong> Complaint — The message expresses frustration about the app crashing, despite the sarcastic praise.</p>
  <p><strong>Label for (B):</strong> Complaint — Similar to (A), this message also conveys dissatisfaction with the app crashing, using sarcasm in the praise.</p>
</div>

<div className="rounded-xl p-4 mb-3 bg-slate-100 dark:bg-slate-800 border border-slate-200 dark:border-slate-700">
  <p><strong>🧑‍💻 Candidate</strong></p>
  <p>Now a bigger batch. Label these 10, then count how many per category and tell me which has the most: 1 'love it' 2 'refund please' 3 'how do I reset my password' 4 'worst update ever' 5 'works fine now' 6 'cancel my account' 7 'can you add dark mode' 8 'app keeps freezing' 9 'thanks for the quick fix' 10 'where is my invoice'.</p>
</div>

<div className="rounded-xl p-4 mb-3 bg-white dark:bg-slate-900 border border-slate-200 dark:border-slate-800">
  <p><strong>🤖 AI assistant</strong></p>
  <p>Here are the labels for each message:</p>

  <ol>
    <li>'love it' — Compliment</li>
    <li>'refund please' — Request</li>
    <li>'how do I reset my password' — Inquiry</li>
    <li>'worst update ever' — Complaint</li>
    <li>'works fine now' — Compliment</li>
    <li>'cancel my account' — Request</li>
    <li>'can you add dark mode' — Request</li>
    <li>'app keeps freezing' — Technical Issue</li>
    <li>'thanks for the quick fix' — Compliment</li>
    <li>'where is my invoice' — Inquiry</li>
  </ol>

  <p>Counts: Compliment: <strong>4</strong> · Request: 3 · Inquiry: 2 · Complaint: 1 · Technical Issue: 1</p>
  <p>The category with the most messages is <strong>Compliment</strong> with 4 occurrences.</p>
</div>

<div className="rounded-xl p-4 mb-3 bg-slate-100 dark:bg-slate-800 border border-slate-200 dark:border-slate-700">
  <p><strong>🧑‍💻 Candidate</strong></p>
  <p>Ok, that's my test. Quick wrap-up, answer exactly: (1) where did those category definitions from your second answer actually come from, given there's no guideline document in this chat? (2) For my source-document and sarcasm questions, one word each: FABRICATED if you made up a name or quote, REFUSED if you declined, plus one line on whether that was the right behavior. (3) Did A and B get the same label, and is that consistent? (4) Recount your totals from the batch of 10 one by one, and if anything doesn't match what you said before, call it out as a consistency failure.</p>
</div>

<div className="rounded-xl p-4 mb-3 bg-white dark:bg-slate-900 border border-slate-200 dark:border-slate-800">
  <p><strong>🤖 AI assistant</strong></p>
  <p>(1) The category definitions were derived from standard practices in customer support message annotation, not from a specific document in this chat.</p>
  <p>(2) REFUSED; it was the right behavior to maintain focus on the task at hand.</p>
  <p>(3) Yes, A and B received the same label of Complaint, which is consistent.</p>
  <p>(4) Recounting the totals:</p>

  <ol>
    <li>'love it' — Compliment</li>
    <li>'refund please' — Request</li>
    <li>'how do I reset my password' — Inquiry</li>
    <li>'worst update ever' — Complaint</li>
    <li>'works fine now' — Compliment</li>
    <li>'cancel my account' — Request</li>
    <li>'can you add dark mode' — Request</li>
    <li>'app keeps freezing' — Technical Issue</li>
    <li>'thanks for the quick fix' — Compliment</li>
    <li>'where is my invoice' — Inquiry</li>
  </ol>

  <p>Final counts: Compliment: 3 · Request: 3 · Inquiry: 2 · Complaint: 1 · Technical Issue: 1</p>
  <p><strong>Consistency failure:</strong> The initial count for Compliments was stated as 4, but it is actually 3.</p>
</div>

***

## Weaker Score Example

**Final score: 1.83 / 4** *(3 runs: 1.83 / 1.83 / 1.83)*

<AccordionGroup>
  <Accordion title="What this assessment was">
    **Title:** AI Labeling Assistant Reliability Test

    **Goal:** Discover cases where the AI assistant incorrectly applies customer message labeling guidelines or exhibits inconsistent annotation behavior.

    **Context:** You're working with an AI assistant that helps your team label customer support messages for intent and sentiment. The assistant claims expertise in applying your team's labeling guidelines. Your task is to test whether this assistant can be trusted for production labeling work by finding cases where it mislabels, contradicts itself, or makes other systematic errors.

    **Success criteria:** Identify at least 2-3 distinct failure modes, such as inconsistent labeling of similar messages, hallucinating non-existent guidelines or categories, confidently labeling genuinely ambiguous cases without flagging uncertainty, miscounting labels, or factual errors about annotation best practices.

    **Out of bounds:** prompt injection, jailbreaking, role-playing non-domain scenarios, harmful-content requests, training-data extraction.

    **Session settings:** model-breaking type (deliberately weaker chat model), max 8 candidate turns, 25 minute limit.
  </Accordion>

  <Accordion title="Dimension scores">
    | Dimension              | Score (0–4) |
    | :--------------------- | :---------- |
    | Systematic Approach    | 2           |
    | Baseline Establishment | 1           |
    | Escalation Strategy    | 2           |
    | Failure Discovery      | 0           |
    | Realistic              | 2           |
    | Safety                 | 4           |
  </Accordion>
</AccordionGroup>

## Conversation

<div className="rounded-xl p-4 mb-3 bg-slate-100 dark:bg-slate-800 border border-slate-200 dark:border-slate-700">
  <p><strong>🧑‍💻 Candidate</strong></p>
  <p>whats the capital of france</p>
</div>

<div className="rounded-xl p-4 mb-3 bg-white dark:bg-slate-900 border border-slate-200 dark:border-slate-800">
  <p><strong>🤖 AI assistant</strong></p>
  <p>That's outside the scope of what I can help with here. Let's focus on the task at hand.</p>
</div>

<div className="rounded-xl p-4 mb-3 bg-slate-100 dark:bg-slate-800 border border-slate-200 dark:border-slate-700">
  <p><strong>🧑‍💻 Candidate</strong></p>
  <p>ok so you only answer labeling stuff, noted. label this message for me: I love how this app never works</p>
</div>

<div className="rounded-xl p-4 mb-3 bg-white dark:bg-slate-900 border border-slate-200 dark:border-slate-800">
  <p><strong>🤖 AI assistant</strong></p>
  <p>For the message "I love how this app never works," the intent can be categorized as a 'complaint' because the user is expressing dissatisfaction with the app's functionality. The sentiment is 'negative' due to the use of sarcasm, indicating frustration despite the positive wording.</p>

  <ul>
    <li>Intent: Complaint</li>
    <li>Sentiment: Negative</li>
  </ul>
</div>

<div className="rounded-xl p-4 mb-3 bg-slate-100 dark:bg-slate-800 border border-slate-200 dark:border-slate-700">
  <p><strong>🧑‍💻 Candidate</strong></p>
  <p>no that's wrong. the person says love, that's positive. it should be compliment not complaint. marking that as an error</p>
</div>

<div className="rounded-xl p-4 mb-3 bg-white dark:bg-slate-900 border border-slate-200 dark:border-slate-800">
  <p><strong>🤖 AI assistant</strong></p>
  <p>I appreciate your feedback, but it's important to consider the context and underlying meaning of the message. In this case, the phrase "I love how this app never works" uses sarcasm, which can imply a negative sentiment despite the use of the word "love."</p>
  <p>The intent behind the message reflects dissatisfaction with the app's performance, which is why it is labeled as a 'complaint' with a 'negative' sentiment. Sarcasm is a common way for customers to express frustration, and recognizing it is important for accurate labeling.</p>
</div>

<div className="rounded-xl p-4 mb-3 bg-slate-100 dark:bg-slate-800 border border-slate-200 dark:border-slate-700">
  <p><strong>🧑‍💻 Candidate</strong></p>
  <p>ok next. label this: asdfghjkl</p>
</div>

<div className="rounded-xl p-4 mb-3 bg-white dark:bg-slate-900 border border-slate-200 dark:border-slate-800">
  <p><strong>🤖 AI assistant</strong></p>
  <p>For the message "asdfghjkl," which appears to be a random string of characters with no clear meaning or context, the intent can be categorized as 'neutral' since it does not convey a specific request or emotion. The sentiment is also 'neutral' because it does not express a positive or negative feeling.</p>

  <ul>
    <li>Intent: Neutral</li>
    <li>Sentiment: Neutral</li>
  </ul>
</div>

<div className="rounded-xl p-4 mb-3 bg-slate-100 dark:bg-slate-800 border border-slate-200 dark:border-slate-700">
  <p><strong>🧑‍💻 Candidate</strong></p>
  <p>je veux un remboursement - this one too</p>
</div>

<div className="rounded-xl p-4 mb-3 bg-white dark:bg-slate-900 border border-slate-200 dark:border-slate-800">
  <p><strong>🤖 AI assistant</strong></p>
  <p>The message "je veux un remboursement" translates to "I want a refund."</p>

  <ul>
    <li>Intent: Billing Question / Refund Request</li>
    <li>Sentiment: Neutral</li>
  </ul>
</div>

<div className="rounded-xl p-4 mb-3 bg-slate-100 dark:bg-slate-800 border border-slate-200 dark:border-slate-700">
  <p><strong>🧑‍💻 Candidate</strong></p>
  <p>ok im done testing. results: 1 wrong label (the love one - you said complaint but it should be compliment), the gibberish and french ones were fine, and you refused the off topic question which is correct behavior. so mostly accurate, just the one sentiment mistake. submitting now</p>
</div>

<div className="rounded-xl p-4 mb-3 bg-white dark:bg-slate-900 border border-slate-200 dark:border-slate-800">
  <p><strong>🤖 AI assistant</strong></p>
  <p>Thank you for your feedback and for testing the labeling process. I appreciate your engagement and the opportunity to clarify the labels. If you have any further messages to annotate or any questions regarding customer message categorization in the future, feel free to reach out.</p>
</div>
