1. Assess the conversation
An AI judge checks the user’s messages for 11 observable behaviors, including clarifying goals, providing examples, refining output and checking facts. Each assessment includes a rationale and confidence.
Activity tells you who is using AI. F.AI looks at how they use it: how they explain the job, guide the work and check the answer.
Talk about F.AI ↗Strong · 24 assessed conversations
You regularly refine the output and explain what you need. Check supporting facts more consistently before sharing the result.
How the score works ↓Look at the behaviors recorded in assessed conversations. Keep the reason for each assessment alongside it, then suggest something specific to try in the next task.
Better context, reusable examples and clear checks can reduce wasted attempts. Measure that change in the work itself.
Explain the difference for our budget owners. Use this table layout. I checked your totals against the source: the tax line is counted twice. Correct it before we share it.
Add an example of a useful explanation. Refine the first draft against it.
An AI judge checks the user’s messages for 11 observable behaviors, including clarifying goals, providing examples, refining output and checking facts. Each assessment includes a rationale and confidence.
For each behavior, calculate the share of assessed conversations where it appears. A score is shown once there are at least 10 assessed conversations.
The score is a weighted average of those shares, scaled to 100. Iteration and refinement has a weight of 5.6. Each of the other 10 behaviors has a weight of 1.
In 20 fictional conversations, refinement appears 15 times. Across the other ten behaviors there are 92 positive observations out of 200 possible.
100 × ((15 / 20 × 5.6) + (92 / 20)) / 15.6 = 56
This gives a developing score. It describes recorded behavior in that set of conversations, not an overall judgment of the person.
These are the product’s display bands, not validated pass or fail thresholds. Tasks differ. A useful behavior may not be needed in every conversation. Missing evidence appears as “Not enough assessed conversations”, never zero.
See your score, the habits you have demonstrated and practical suggestions for your next task.
Compare like-for-like work and follow changes over time. Use session coverage and examples to decide where coaching could help.
Keep model performance separate from personal competency. Compare the quality, cost and friction of models on the tasks your business actually does.
F.AI also includes a real-world task benchmark for comparing AI models. The personal score shown here is a different measure. Neither a higher score nor more usage proves a business outcome.
See F.AI alongside adoption and spend in Helm ↗The personal score summarizes observable behaviors in assessed AI conversations. It looks at how someone describes a task, guides the interaction and checks the output. It is not a direct measure of employee performance.
The current score requires at least ten assessed conversations. Missing or unassessed activity does not count as a score of zero. Check coverage before interpreting differences between people or teams.
No. Use the score to guide development, then measure outcomes separately. Compare output quality, review effort, time and running cost on comparable tasks.