Research

We build benchmarks for the decisions LearnOS actually asks a model to make, then publish the method and the per-model numbers — including the ones that say nothing separates.

Category Bv0.1.0 · 150 tasks

Tool-use routing

Given a learner's message, which of twelve tools should fire — and with which arguments? Scores the decision only; nothing is dispatched.

Category APlanned

Grounded tutoring

Citation validity, gold-anchor attribution, and abstention when the corpus has no answer. Most of it is machine-checkable.

Category CPlanned

Course authoring

Write a lesson from a topic and a source list, then score it on the same six-category rubric we run against imported courses.

Category DPlanned

Structured generation

Presentations, audio overviews, syllabi — schema conformance and constraint adherence.

Latest run 2026-09-19. Method and task file are open — see the Category B results.

Command Palette

Search for a command to run...