Skip to main content

Evals: test your AI’s answers

Use an eval to send a dataset of examples through one AI node and see how many of its answers passed.

An eval tests one AI node in a workflow. It sends each example in its dataset to that node, and a judge scores each answer. Then you see a simple result, like 8 of 10 passed.

Run an eval before you change a prompt or a model, and again after. If the number drops, the change made your AI worse.

Evals live under AI → Evals in your workspace sidebar.

The words you will see

WordWhat it means
EvalA test for one AI node. It has a name, the node it tests, and a dataset.
DatasetThe examples an eval uses. Each example is a question your AI node should answer well.
JudgeAn AI model that reads each answer and gives it a score. It uses your own AI key.
RubricThe judge instructions: what the judge, or a person, follows to decide if an answer is good.
RunOne pass of an eval over its dataset.
ResultWhat a run tells you, for example 6 of 10 answers passed.
ReviewA list of answers for a person to score by hand, to make sure the judge agrees with real people.

Make your first eval

Go to AI → Evals and click New eval. It takes about 5 minutes. You can also start from the workflow editor: open the menu on an AI node and choose Evals.

  1. Pick the AI node

    Pick the AI node you want to evaluate. Nodes are listed by workflow.

  2. Add examples

    Add at least 5. Real questions from your customers work best. You can:

    • From recent runs: tick runs where the node did its job. Each one becomes an example. Private details like names and emails are taken out first.
    • Type them in: write a question, and what a good answer looks like if you want.
    • Upload a spreadsheet: use a CSV file with a question column and an optional good_answer column. Download the template from the same screen.
    • Let AI write some: handy for trying an eval, but these examples don't count toward pass or fail.
  3. Choose how to score answers

    Pick one:

    • Should match the good answer: the judge compares each answer to the good answer you gave.
    • Should follow these instructions: you write the rubric, and the judge checks each answer against it.
    • Should pick the right label: for a node that sorts things into groups. The label must be the same as the good answer.

    Then pick a pass mark. At 70%, an answer passes when it scores 70% or higher, and the eval passes when at least 70% of answers pass.

  4. Confirm and run

    Look over your eval, then click Run eval. You land on the result. It fills in as the judge scores each answer.

To add more examples later, open the eval and click See or add examples. The eval keeps using the newest version of its dataset.

When an eval passes

Each answer gets a score from the judge. An eval passes when enough of its answers pass. New evals start at 70%, so with 10 examples, at least 7 answers must pass.

Evals are advice. A failed eval does not stop you from publishing a workflow.

What an eval costs

Each run asks the judge to score every example once. The judge uses your own AI key, so the cost goes on your AI provider's bill, not on TaskJuice. The eval's page shows a rough cost per run. Open Settings on that page to see the math, or to pick a different key or model for the judge.

If you can't run an eval

The Run eval button is turned off and the page says why. The common reasons:

What the page saysWhat to do
You need at least one exampleAdd a real example. Examples made by AI help you try an eval, but they don't count.
The judge has no AI key yetOpen Settings on the eval's page and pick an AI key and model.
This eval uses an old version of its datasetMake a new eval. It will use the latest examples.
Was this helpful?