The new scoreboard. Four layers, one system
This essay unveils a layered system measuring learner engagement and mastery, crucial for identifying and addressing educational barriers effectively.
1) The new scoreboard. Four layers, one system
Layer A. Access and Engagement. “Can the learner enter the work
” This is where meeting kids where they are lives. It is not a product nicety. It is a gating constraint on durable mastery. Layer B. Mastery Integrity. “Is mastery real, durable, and transferable” This is your current center of gravity. Retention, transfer, false mastery, calibration, mastery depth. Layer C. One-Degree Control. “Does the system collapse distance between evidence and action” This is the disruptive layer. It turns learning into controllability. Layer D. Outcomes. “What stakeholders recognize and fund” This is where you bridge to reading improvement, course mastery, benchmark movement, and teacher capacity. Each layer has KPIs that are computable from event logs. Each KPI has a “why it matters,” and a direct lever in the platform.
2) Layer A KPIs. Access and Engagement. The missing half of the system
A1. Concept Access Friction
Definition: expected steps lost before a learner can produce valid evidence on a concept due to language load, not concept load. Compute: time to first valid paraphrase plus clarifications plus abandoned attempts, normalized by concept difficulty. Why this matters: your traversal metrics assume the learner can engage. Unengaged learners are often “access blocked,” not “ability blocked.” A2. Vocabulary Coverage Ratio Definition: percent of concept critical terms the learner can correctly use in context during explanation. Not definition recall. Use in context. Compute: term usage correctness within explanations and paraphrases. Why this matters: vocabulary is the hidden prerequisite graph. If you do not model it, you misroute instruction. A3. Reading Load Delta Definition: the gap between the reading complexity of prompts and the learner’s demonstrated reading capacity. This is “meeting them where they are” as a number. Compute: readability and sentence load of prompt, minus learner reading band inferred from paraphrase accuracy, comprehension checks, and error signatures. Why this matters: if the delta is too high, you get false mastery signals, disengagement, and bad causal attribution. A4. Prompt Understanding Rate Definition: percent of tasks where the learner correctly identifies what the question is asking. This is not content mastery. It is entry. Compute: paraphrase pass rate plus constraint recognition rate plus “answered the wrong question” detection. A5. On-Ramp Conversion Rate Definition: percent of near-zero evidence learners who move into “steady evidence production” within two weeks. Compute: band transition based on Evidence Density thresholds. Why this matters: this is the unengaged rescue KPI. If this does not move, nothing scales. A6. Evidence Density Definition: meaningful evidence events per learner per week. Not minutes. Not logins. Evidence events include. paraphrase pass, explanation submitted, misconception corrected, transfer success, vocabulary used correctly, durable recheck pass. A7. Interest Alignment Lift Definition: the incremental gain in Evidence Density and Explanation Quality when the same concept is presented through different context wrappers that align to learner interests. Why this matters: you are testing a causal claim. Meeting kids where they are increases engagement and therefore increases learning opportunity count. This KPI also prevents self deception. It forces you to prove that “interest adaptation” is doing work. A8. Frustration Stall Rate Definition: percent of sessions that terminate after repeated access failures before any valid evidence is produced. Why this matters: stall is the signature of “distance” between the learner and the work. This is the one-degree problem inside the student.
3) Layer B KPIs. Mastery Integrity. Upgrade what you already have
You already track most of these. The upgrade is to split them into “real mastery” vs “fragile mastery,” and to connect them to the access layer. B1. Mastery Depth Keep your definition. Depth is relational understanding across prerequisites and downstream concepts. Upgrade. compute depth as a graph property, not a per item property. Add: Depth Consistency Score How stable the depth estimate is across different prompt contexts, reading levels, and surface forms. B2. Mastery Convergence Steps Keep. steps to stable mastery threshold. Upgrade. report convergence as a distribution, not a scalar. Median and tail. Add: Tail Convergence Penalty The system is only as good as its slowest cohort. This is the equity and scalability lever. B3. Retention Keep. but make it operational. Add: Forgetting Slope Rate of decay in mastery probability over time windows. Add: Retention Under Context Shift Retention when the concept is presented with different vocabulary, different framing, or different domain examples. B4. Transfer Keep. but disambiguate. Add: Near Transfer vs Far Transfer Near transfer. new surface form, same structure. Far transfer. different domain, same causal structure. B5. False Mastery Keep. but attribute it. Add: False Mastery Cause Codes Reading load caused, vocabulary caused, lucky guessing caused, rubric mismatch caused, causal model gap caused. This is essential. Otherwise you treat all false mastery as model failure. B6. Calibration Error Keep. but report two forms. Overconfidence Rate: high confidence on wrong. This is the safety risk. Underconfidence Rate: low confidence on right. This is the efficiency loss. B7. Learning Efficiency Keep, but upgrade it. Durable Learning per Step: mastery probability gain that survives retention checks. Causal Gain per Step: expected reduction in uncertainty about the learner’s misconception fingerprint, not just correctness. B8. Cost per Durable Mastery Keep. It is your CFO metric. Upgrade. include rework cost from false mastery and access friction.
4) Layer C KPIs. One-Degree Control. What makes this disruptive
This is the layer that turns Trek into an operating system. It measures distance collapse. C1. Learning Latency Definition: time from first evidence of misunderstanding to verified state change. You already know this one. Make it the headline. Break into sub latencies so you can attack bottlenecks. • Detect latency • Diagnose latency • Route latency • Act latency • Verify latency C2. Degree of Separation Index Definition: number of handoffs and permission gates between evidence and action. Compute by counting transitions across roles, tools, approvals, and waiting states. Why this matters: most schools are slow because the architecture is slow. If Trek is one degree away, this index must drop. C3. Loop Closure Rate Definition: percent of detected misconceptions or access blocks that reach verified resolution inside SLA. C4. Action Ownership Integrity Definition: percent of recommended actions that have. named owner, due date, success criteria, escalation rule. If this is not near perfect, you are a suggestion engine. Not a control system. C5. Intervention Fidelity Definition: did the prescribed action actually happen. With the correct dosage, correct target, and correct follow up evidence. C6. Minimal Effective Intervention Definition: smallest action that produces verified mastery movement. This is the causal engine KPI. Compute by testing interventions in increasing intensity order, measuring marginal gain, and stopping when gain crosses threshold. Why this matters: it is how you scale without burning teachers. C7. Counterfactual Accuracy Definition: when the system claims a cause, does the predicted correction work when applied. This is the automated reasoning integrity metric. Example. If the system says the failure is vocabulary, then applying vocabulary scaffolding should produce a predictable shift. If not, causal attribution was wrong. C8. Target Learning Focus You already compute something like this. Upgrade it into a full decision policy KPI. Target Focus Quality: how often the selected next concept is on the highest expected durable gain path given constraints. Access Aware Focus: how often the system selects a concept that the learner can actually enter today. A system that always picks the “optimal concept” but ignores access friction is not optimal. It is naive. C9. Recovery Half-Life Definition: median time to pull a learner out of a failure regime. • out of near zero evidence • out of repeated misconception recurrence • out of chronic prompt misread This is your “resilience” metric.
5) Layer D KPIs. Outcomes that stakeholders will believe
If you want to reset the paradigm, outcomes become a roll up, not the engine. Still, you need them. D1. Reading Ceiling Reduction Definition: reduction in the share of failures attributable to reading and vocabulary friction. If Trek improves reading, this is the mechanism metric that shows how. D2. Cross Subject Mastery Lift Definition: mastery velocity gains in science and math explained by reductions in access friction and increased evidence density. This ties the “meeting them where they are” layer to observable achievement. D3. Durable Mastery Coverage Definition: percent of priority standards with durable mastery, not transient pass. D4. Teacher Time Returned Definition: minutes reduced in grading and diagnosis, and minutes reallocated to targeted instruction. If you cannot show this, adoption dies, regardless of learning gains. D5. Variance Compression Definition: reduction in outcome variance across classrooms and cohorts. This is the “system gets less dependent on hero teachers” KPI. D6. Cost per Verified Mastery Gain Your procurement KPI. If you want to be disruptive, price to this, not to seats.
6) The concept graph itself. New graph metrics you should add
You are evaluating traversal across concepts and steps. Add these. They expose whether the tutor is learning like a human and teaching like a control system. G1. Prerequisite Respect Rate Percent of time the system attempts downstream concepts before prerequisites are truly mastered. This is a major source of false mastery and unengagement. G2. Productive Revisitation Rate Percent of revisits that yield durable gain. Revisits are not bad. Unproductive revisits are. G3. Misconception Path Signature The recurring path patterns that indicate the same wrong mental model propagating across concepts. This is where causal graphs beat naive tutoring. G4. Graph Coherence Score How consistent the learner’s inferred knowledge state is with the prerequisite structure and dependency edges. If the system claims mastery on a node while prerequisite nodes are weak, coherence drops. That is a warning. It is also a trigger for targeted remediation. G5. Access Graph Overlay This is new. Add a second graph that represents language prerequisites. • vocabulary nodes • reading complexity gates • prompt comprehension gates Then compute. concept traversal is only permitted when access overlay gates are satisfied or intentionally scaffolded. This is meeting kids where they are as an architecture. G6. Intervention Routing Optimality When an access failure is detected, how often does the system route to the correct scaffold type. • vocabulary scaffold • paraphrase scaffold • reduced reading load • interest aligned framing • prerequisite refresh • misconception correction This is a measurable policy.
7) Instrumentation. The event schema Trek must log to compute all of this
If you want board grade credibility, every KPI must be computable from auditable events. Evidence events • paraphrase_attempted, paraphrase_passed • prompt_constraint_detected • vocab_term_used_in_context, vocab_term_misused • explanation_submitted, explanation_quality_scored • misconception_detected, misconception_fingerprint_assigned • mastery_probability_updated, uncertainty_updated • transfer_item_presented, transfer_passed • retention_check_scheduled, retention_check_passed Control events • intervention_recommended • intervention_assigned, owner, due_date • intervention_started, intervention_completed, dosage_minutes • verification_evidence_collected, verification_passed, verification_failed • escalation_triggered, escalated_to_role Access and engagement events • reading_load_level_of_prompt • access_scaffold_applied_type • stall_detected_reason • session_abandoned_reason • interest_context_selected • evidence_density_counter_increment With this, the system can be audited. No vanity metrics. No unprovable claims.
8) The stakeholder scorecards. Same engine, different views
Student view
• Mastery state by concept • Next best step. Access scaffold if needed • Proof of progress. durable and transfer, not just a green check Teacher view • Top access blockers. vocabulary, prompt misread, reading load • Highest leverage interventions. minimal effective • Evidence backlog needing action. ordered by latency risk • Time returned. and where to spend it Principal view • Learning latency distribution • Loop closure rate • Near zero band size and recovery half life • Variance compression across classrooms District and board view • Cost per verified mastery gain • Reading ceiling reduction • Durable mastery coverage • Degree of separation index. trend • Governance. audit completeness, policy compliance If the board cannot see controllability improving, they will not fund it. If teachers cannot feel time returning, they will not use it. If students cannot enter the work, nothing else matters.
9) The key move you just implied. Access is part of the causal model
Meeting kids where they are is not a UX feature. It is causal. If the system reduces reading load and scaffolds vocabulary, it should produce: • higher evidence density • lower stall rate • lower false mastery caused by prompt misunderstanding • faster mastery convergence • higher transfer, because the learner can actually practice the concept, not just endure the prompt So you must explicitly model access as a first class causal variable, not as noise. That is how you keep the system honest and why this becomes a true automated reasoning machine.
References
Internal context provided by you in this thread. • Screenshot showing the current evaluation layout. Mastery radar across 53 concepts, learning KPIs including retention, transfer, false mastery, calibration, efficiency, cost per durable, durable count. System metrics including mastery depth, learning realism, convergence steps, and targeted learning focus. • Metric definitions you supplied. concept mastery, mastery convergence, retention, transfer, false mastery, calibration error, learning efficiency, time to first concept access, learning realism, cost per durable mastery, targeted learning focus. Foundational measurement ideas that align with the KPI design. • Bloom. mastery learning and corrective feedback as a structural advantage. • Black and Wiliam. formative assessment as a feedback loop that changes instruction. • Scarborough. reading as a constrained system with vocabulary and language comprehension as limiting factors.