Research
We build benchmarks for the decisions LearnOS actually asks a model to make, then publish the method and the per-model numbers — including the ones that say nothing separates.
Tool-use routing
Given a learner's message, which of twelve tools should fire — and with which arguments? Scores the decision only; nothing is dispatched.
Grounded tutoring
Citation validity, gold-anchor attribution, and abstention when the corpus has no answer. Most of it is machine-checkable.
Course authoring
Write a lesson from a topic and a source list, then score it on the same six-category rubric we run against imported courses.
Structured generation
Presentations, audio overviews, syllabi — schema conformance and constraint adherence.
Latest run 2026-09-19. Method and task file are open — see the Category B results.
