EdTech companies love to celebrate metrics. Minutes spent in the app. Problems completed. Streaks maintained. These are easy to measure and easy to display on a dashboard. But none of them answer the question that matters most to schools, parents, and investors: is the product actually helping students learn?
Why Activity Metrics Are Not Outcomes
Activity metrics tell you whether a student opened the app. They do not tell you whether the student understood the material. A learner can spend thirty minutes in a math tutoring app and get every problem wrong — and the platform would still report thirty minutes of engagement. This is the fundamental problem with measuring EdTech success purely through activity data.
Learning outcomes, on the other hand, measure actual skill acquisition or knowledge transfer. Did test scores improve? Can the student apply the concept in a new context? Did mastery persist over time? These questions are harder to answer, but they are the only questions that matter when evaluating whether an AI learning product delivers value.
A Framework for Measuring What Actually Matters
Nivorius recommends a four-tier framework for measuring learning outcomes in AI-powered EdTech products. Each tier serves a different stakeholder and answers a different question.
Tier 1: Proximal Metrics (Learner-Level)
Proximal metrics measure immediate skill acquisition. In an AI tutoring context, these include diagnostic assessment scores, concept mastery percentages, and the rate of correct responses over time. These metrics tell you whether the AI is adapting effectively to an individual learner's needs.
- Pre- and post-diagnostic scores within a learning session
- Concept mastery progression tracked by the adaptive engine
- Error pattern analysis showing which skill gaps are being addressed
- Time-to-mastery for individual concepts or skills
Tier 2: Distal Metrics (Classroom-Level)
Distal metrics connect product usage to classroom performance. This requires integration with existing assessment systems or curriculum-aligned benchmarks. Schools need to know whether using the product correlates with improved grades, standardized test scores, or teacher-reported performance.
- Correlation between product usage and grade improvements
- Comparison between product users and non-users in the same classroom
- Teacher-reported skill improvements that align with product learning objectives
- Curriculum benchmark alignment scores
Tier 3: Behavioral Metrics (Engagement-Level)
Behavioral metrics capture whether learning is transferring to real-world contexts. Does the student apply concepts outside the app? Do they seek out additional learning opportunities? These metrics are harder to measure but provide crucial evidence of genuine learning versus surface-level engagement.
- Spontaneous use of learned concepts in new contexts
- Transfer of skills to different subjects or platforms
- Self-directed learning behavior outside mandated sessions
- Persistence of engagement over extended time periods
Tier 4: Longitudinal Metrics (System-Level)
Longitudinal metrics track learning impact over extended periods — months or years. This is where AI products prove their long-term value. Do students who used the product in 5th grade show stronger performance in 6th grade? Did the product reduce the need for remediation? These metrics require sustained data collection and are the most expensive to measure, but they are also the most persuasive for district-level decision makers.
- Year-over-year performance trajectory for product users
- Remediation rate reduction among persistent users
- College readiness indicators for sustained product users
- Cross-cohort comparison of long-term learning outcomes
What Most EdTech Products Get Wrong
The most common mistake is confusing Tier 1 metrics with Tier 2 or Tier 3 outcomes. When a product claims that students who use the platform show '40% improvement,' the fine print usually reveals that this improvement is measured through in-app diagnostics — not through external assessments or real-world performance.
Another common error is measuring short-term gains without tracking whether those gains persist. An AI tutoring product might help a student pass a unit test, but if the knowledge evaporates within weeks, the product is not delivering lasting educational value.
Building a Measurement Framework That Scales
For early-stage EdTech products, start with Tier 1 metrics and build from there. Every product should be able to answer whether its AI is helping individual learners acquire skills. From there, integrate with classroom assessments to track Tier 2 metrics. Tier 3 and Tier 4 require partnerships with schools and districts but become powerful differentiators for products aiming at enterprise adoption.
The products that will win in EdTech are not the ones with the flashiest dashboards. They are the ones that can demonstrate — with credible data — that their AI is making students better learners. That starts with measuring the right things.
Part of the Nivorius research and consulting team, focused on practical applications of AI in education and enterprise contexts.

