- Documentation
- Workflows
- Evals: test your AI’s answers
Evals: test your AI’s answers
Use an eval to send a dataset of examples through one AI node and see how many of its answers passed.
An eval tests one AI node in a workflow. It sends each example in its dataset to that node, and a judge scores each answer. Then you see a simple result, like 8 of 10 passed.
Run an eval before you change a prompt or a model, and again after. If the number drops, the change made your AI worse.
Evals live under AI → Evals in your workspace sidebar.
The words you will see
| Word | What it means |
|---|---|
| Eval | A test for one AI node. It has a name, the node it tests, and a dataset. |
| Dataset | The examples an eval uses. Each example is a question your AI node should answer well. |
| Judge | An AI model that reads each answer and gives it a score. It uses your own AI key. |
| Rubric | The judge instructions: what the judge, or a person, follows to decide if an answer is good. |
| Run | One pass of an eval over its dataset. |
| Result | What a run tells you, for example 6 of 10 answers passed. |
| Review | A list of answers for a person to score by hand, to make sure the judge agrees with real people. |
Make your first eval
Go to AI → Evals and click New eval. It takes about 5 minutes. You can also start from the workflow editor: open the menu on an AI node and choose Evals.
Pick the AI node
Pick the AI node you want to evaluate. Nodes are listed by workflow.
Add examples
Add at least 5. Real questions from your customers work best. You can:
- From recent runs: tick runs where the node did its job. Each one becomes an example. Private details like names and emails are taken out first.
- Type them in: write a question, and what a good answer looks like if you want.
- Upload a spreadsheet: use a CSV file with a
questioncolumn and an optionalgood_answercolumn. Download the template from the same screen. - Let AI write some: handy for trying an eval, but these examples don't count toward pass or fail.
Choose how to score answers
Pick one:
- Should match the good answer: the judge compares each answer to the good answer you gave.
- Should follow these instructions: you write the rubric, and the judge checks each answer against it.
- Should pick the right label: for a node that sorts things into groups. The label must be the same as the good answer.
Then pick a pass mark. At 70%, an answer passes when it scores 70% or higher, and the eval passes when at least 70% of answers pass.
Confirm and run
Look over your eval, then click Run eval. You land on the result. It fills in as the judge scores each answer.
To add more examples later, open the eval and click See or add examples. The eval keeps using the newest version of its dataset.
When an eval passes
Each answer gets a score from the judge. An eval passes when enough of its answers pass. New evals start at 70%, so with 10 examples, at least 7 answers must pass.
Evals are advice. A failed eval does not stop you from publishing a workflow.
What an eval costs
Each run asks the judge to score every example once. The judge uses your own AI key, so the cost goes on your AI provider's bill, not on TaskJuice. The eval's page shows a rough cost per run. Open Settings on that page to see the math, or to pick a different key or model for the judge.
If you can't run an eval
The Run eval button is turned off and the page says why. The common reasons:
| What the page says | What to do |
|---|---|
| You need at least one example | Add a real example. Examples made by AI help you try an eval, but they don't count. |
| The judge has no AI key yet | Open Settings on the eval's page and pick an AI key and model. |
| This eval uses an old version of its dataset | Make a new eval. It will use the latest examples. |