Bayesian Knowledge Tracing (BKT) is a probabilistic modelling technique that estimates a learner’s mastery of a skill over time by treating knowledge as a latent binary state inferred from a sequence of correct and incorrect responses. Using a hidden Markov model with parameters for prior knowledge, learning, guessing, and slipping, BKT updates the probability that a student has mastered each skill after every interaction. It is a cornerstone of intelligent tutoring systems and adaptive learning, enabling personalised pacing and content selection.
Semantic Classification
Content
Compositional Relationships (Components)
SubClassOf(edu:BayesianKnowledgeTracing
ObjectSomeValuesFrom(edu:hasPart edu:HiddenMarkovModel))
SubClassOf(edu:BayesianKnowledgeTracing
ObjectSomeValuesFrom(edu:hasPart edu:KnowledgeComponentModel))
SubClassOf(edu:BayesianKnowledgeTracing
ObjectSomeValuesFrom(edu:hasPart edu:ProbabilisticInference))
SubClassOf(edu:BayesianKnowledgeTracing
ObjectSomeValuesFrom(edu:hasPart edu:MasteryThreshold))
SubClassOf(edu:BayesianKnowledgeTracing
ObjectSomeValuesFrom(edu:hasPart edu:ParameterEstimation))
SubClassOf(edu:BayesianKnowledgeTracing
ObjectSomeValuesFrom(edu:hasPart edu:BayesianInference))
Dependency Relationships
SubClassOf(edu:BayesianKnowledgeTracing
ObjectSomeValuesFrom(edu:requires edu:KnowledgeComponentModel))
SubClassOf(edu:BayesianKnowledgeTracing
ObjectSomeValuesFrom(edu:requires edu:EducationalTechnology))
SubClassOf(edu:BayesianKnowledgeTracing
ObjectSomeValuesFrom(edu:requires edu:FormativeAssessment))
SubClassOf(edu:BayesianKnowledgeTracing
ObjectSomeValuesFrom(edu:dependsOn edu:HiddenMarkovModel))
SubClassOf(edu:BayesianKnowledgeTracing
ObjectSomeValuesFrom(edu:dependsOn edu:BayesianInference))
SubClassOf(edu:BayesianKnowledgeTracing
ObjectSomeValuesFrom(edu:dependsOn edu:ExpectationMaximisation))
SubClassOf(edu:BayesianKnowledgeTracing
ObjectSomeValuesFrom(edu:dependsOn edu:Psychometrics))
Capability Relationships
SubClassOf(edu:BayesianKnowledgeTracing
ObjectSomeValuesFrom(edu:enables edu:PersonalisedLearning))
SubClassOf(edu:BayesianKnowledgeTracing
ObjectSomeValuesFrom(edu:enables edu:MasteryLearning))
SubClassOf(edu:BayesianKnowledgeTracing
ObjectSomeValuesFrom(edu:enables edu:FormativeAssessment))
SubClassOf(edu:BayesianKnowledgeTracing
ObjectSomeValuesFrom(edu:enables edu:LearningAnalytics))
SubClassOf(edu:BayesianKnowledgeTracing
ObjectSomeValuesFrom(edu:enables edu:ComputerisedAdaptiveTesting))
SubClassOf(edu:BayesianKnowledgeTracing
ObjectSomeValuesFrom(edu:supports edu:ComputerisedAdaptiveTesting))
SubClassOf(edu:BayesianKnowledgeTracing
ObjectSomeValuesFrom(edu:supports edu:AdaptiveLearning))
SubClassOf(edu:BayesianKnowledgeTracing
ObjectSomeValuesFrom(edu:supports edu:OpenLearnerModel))
Implementation Relationships
SubClassOf(edu:BayesianKnowledgeTracing
ObjectSomeValuesFrom(edu:implements edu:AdaptiveLearning))
SubClassOf(edu:BayesianKnowledgeTracing
ObjectSomeValuesFrom(edu:implements edu:MasteryLearning))
SubClassOf(edu:BayesianKnowledgeTracing
ObjectSomeValuesFrom(edu:implements edu:BayesianInference))
SubClassOf(edu:BayesianKnowledgeTracing
ObjectSomeValuesFrom(edu:uses edu:HiddenMarkovModel))
SubClassOf(edu:BayesianKnowledgeTracing
ObjectSomeValuesFrom(edu:uses edu:ExpectationMaximisation))
SubClassOf(edu:BayesianKnowledgeTracing
ObjectSomeValuesFrom(edu:uses edu:KnowledgeComponentModel))
Reduction Relationships
SubClassOf(edu:BayesianKnowledgeTracing
ObjectSomeValuesFrom(edu:reducesTo edu:TwoStateProbabilisticModel))
SubClassOf(edu:BayesianKnowledgeTracing
ObjectSomeValuesFrom(edu:relatedTo edu:ItemResponseTheory))
SubClassOf(edu:BayesianKnowledgeTracing
ObjectSomeValuesFrom(edu:relatedTo edu:DeepKnowledgeTracing))
SubClassOf(edu:BayesianKnowledgeTracing
ObjectSomeValuesFrom(edu:relatedTo edu:SpacedRepetition))
SubClassOf(edu:BayesianKnowledgeTracing
ObjectSomeValuesFrom(edu:relatedTo edu:CognitiveTutor))
SubClassOf(edu:BayesianKnowledgeTracing
ObjectSomeValuesFrom(edu:bridgesTo edu:ExplainableAI))
SubClassOf(edu:BayesianKnowledgeTracing
ObjectSomeValuesFrom(edu:bridgesTo edu:DeepKnowledgeTracing))
SubClassOf(edu:BayesianKnowledgeTracing
ObjectSomeValuesFrom(edu:bridgesTo edu:LargeLanguageModel))
About
Bayesian Knowledge Tracing stands as one of the most enduring and practically deployed algorithms at the intersection of Cognitive Science, Psychometrics, and Machine Learning. Its origins lie in the cognitive tutoring research programme led by John R. Anderson at Carnegie Mellon University in the late 1980s and early 1990s, which sought to create computational models of human skill acquisition grounded in Anderson’s ACT-R (Adaptive Control of Thought—Rational) cognitive architecture. Albert Corbett and John Anderson formalised BKT in a 1994 technical report and a canonical 1995 journal paper in User Modeling and User-Adapted Interaction, presenting it as the mastery-tracking component of the CMU Cognitive Tutor for algebra — a system that went on to be deployed in hundreds of thousands of US high school classrooms. The insight was elegant: rather than treating a student’s knowledge state as directly observable (as in simple right/wrong tallies), or as a continuous latent variable requiring expensive psychometric testing (as in Item Response Theory), BKT treats knowledge acquisition as a two-state latent Markov process. The learner occupies either the “not mastered” state or the “mastered” state, and Bayesian updating after each observed response maintains a calibrated posterior probability over which state the learner currently occupies. This posterior then drives the adaptive tutoring decision: practise more (if P(L_t) < threshold) or advance (if P(L_t) ≥ threshold). The mathematical machinery is deliberately minimal: four parameters fully characterise a skill’s learning dynamics, these parameters carry direct pedagogical meaning (prior knowledge, learning rate, guessing tendency, slipping tendency), and the update equations are analytically tractable without any approximation, GPU, or deep learning infrastructure required.
The model’s durability in the face of more powerful neural competitors is explained by four interconnected factors. First, interpretability: BKT’s parameters have direct cognitive and pedagogical meaning, making it possible for curriculum designers and teachers to understand and validate the model’s behaviour in ways that LSTM or Transformer Architecture-based Learner Models do not permit; a product manager can interrogate why a student was held on a skill and receive an intelligible answer in terms of four parameters rather than 175 million attention weights. Second, data efficiency: BKT requires only dozens of observations per skill to converge on reasonable parameter estimates — critical in educational contexts where rare curriculum skills may appear only a handful of times per student — while deep learning approaches require hundreds or thousands of observations to achieve stable predictions. Third, regulatory compliance: Explainable AI requirements under GDPR Article 22 and the UK Department for Education’s 2024 guidance on automated educational decisions favour models whose recommendations can be traced to interpretable parameters and explained to learners and parents. Fourth, integration with Mastery Learning-based pedagogy: BKT directly operationalises Benjamin Bloom’s mastery learning vision — no student advances until mastery is demonstrated — in a way that is structurally natural for the fixed-threshold BKT mastery criterion, and less so for continuous neural predictions that lack a principled natural threshold. These advantages have kept BKT in active production use at Carnegie Learning (MATHia), Duolingo (hybrid implementations with Spaced Repetition), and ASSISTments, even as research benchmarks increasingly favour its deep learning successors. The platform-scale evidence is substantial: Carnegie Learning’s MATHia system, which uses BKT as its mastery model for ~500 algebra and geometry knowledge components, reports learning gains consistent with Bloom’s Two-Sigma Problem predictions when students engage for 30+ minutes per week.
Formal Model and Algorithm
The BKT model is a two-state first-order Hidden Markov Model in which the latent state L_t ∈ {0, 1} (not mastered, mastered) evolves as a Markov chain and generates the observable response O_t ∈ {correct, incorrect} according to emission probabilities parameterised by the guess and slip rates. Formally, for each knowledge component k:
State transition: P(L_t = 1 | L_{t-1} = 0) = P(T) (learn); P(L_t = 1 | L_{t-1} = 1) = 1 (no forgetting in standard BKT, reflecting within-session retention).
Emission probabilities: P(O_t = correct | L_t = 1) = 1 − P(S); P(O_t = correct | L_t = 0) = P(G).
Update equations after observing O_t via Bayesian Inference:
Stage 1 — update on observed response (Bayesian filtering):
-
If O_t = correct: P(L_t | O_t) = [P(L_{t−1}) × (1 − P(S))] / [P(L_{t−1}) × (1 − P(S)) + (1 − P(L_{t−1})) × P(G)]
-
If O_t = incorrect: P(L_t | O_t) = [P(L_{t−1}) × P(S)] / [P(L_{t−1}) × P(S) + (1 − P(L_{t−1})) × (1 − P(G))]
Stage 2 — propagate through the learning transition:
-
P(L_{t+1}) = P(L_t | O_t) + (1 − P(L_t | O_t)) × P(T)
Mastery decision: Declare mastery when P(L_t) ≥ θ, where θ = 0.95 is the conventional threshold used in Carnegie Learning deployments, corresponding to a 95% posterior probability of mastery.
Parameter estimation: Parameters {P(L₀), P(T), P(G), P(S)} are estimated from historical response logs using the Baum-Welch algorithm — a special case of Expectation-Maximisation for Hidden Markov Models. The E-step computes expected state occupancy counts; the M-step updates parameters to maximise expected log-likelihood. Constraints P(G) + P(S) < 1 (degenerate if not enforced) and P(T) ∈ (0, 1) are applied. Baker et al. (2008) documented model identifiability problems where multiple parameter settings produce identical likelihoods on training data, leading to contextual slip estimation extensions and grid search with degeneracy checks as standard practice.
Extensions and Variants
-
Individualised BKT (Yudelson et al., 2013): fits separate P(L₀) and P(T) for each student–skill combination, treating student ability as a modifier on population-level parameters. Demonstrated on the KDD 2010 Algebra dataset, achieving 3–5% AUC improvement over standard BKT at the cost of requiring sufficient per-student data — typically 20+ observations per skill per student — for stable estimation.
-
KT-IDEM (Pardos & Heffernan, 2011): Item-Difficulty Effect Model separates slip and guess parameters by item, allowing items of varying difficulty within the same skill to have different emission probabilities — a bridge between BKT and Item Response Theory that improves performance on heterogeneous item banks.
-
BKT+Forget (Qiu et al., 2011): introduces a non-zero forgetting parameter P(F) = P(L_t = 0 | L_{t−1} = 1) to capture skill decay between practice sessions, important for longitudinal tracking across days or weeks where Spaced Repetition scheduling interacts with knowledge retention.
-
BKT-LSTM hybrid: augments an LSTM predictor with the BKT posterior P(L_t) as an explicit input feature alongside raw response history embeddings, yielding interpretable BKT-calibrated predictions with the sequential representation learning of deep networks. Performance comparable to Deep Knowledge Tracing on ASSISTments benchmarks with enhanced interpretability.
-
B²KT — Bayesian-Bayesian Knowledge Tracing (van der Graaf et al., 2022): treats the per-student BKT parameter vector itself as a random variable with a hierarchical prior, enabling online Bayesian updating of both the knowledge state and the skill parameters simultaneously. Published at EDM 2022, this approach was shown to be more equitable across learners of varying reading ability and demographic groups than classical BKT, directly addressing the fairness concerns dominant in contemporary BKT literature. EDM 2025 presented follow-on work specifically on BKT fairness for math learners of different reading ability.
-
Neural-Symbolic BKT (arXiv:2604.08263, 2026): injects Knowledge Graph structure from curriculum prerequisite graphs into BKT-derived learner models using Graph Neural Network priors, enabling parameter transfer from well-observed to data-sparse skills and combining BKT’s interpretability with the representational power of symbolic knowledge representations.
-
Sparse Binary Representation Learning for KT (arXiv:2501.09893, 2025): introduces sparse binary latent codes alongside response prediction, improving interpretability and enabling compression of knowledge state representations for deployment on resource-constrained tutoring devices.
-
LLM-BKT integration (2024–2026): emerging architectures use Large Language Model backends for dialogue tutoring and explanation generation, with BKT handling the structured mastery-state estimation and pacing decisions — an explicit division of labour that plays to both components’ strengths, deployed in platforms such as Carnegie Learning’s MathGPT and Khan Academy’s Khanmigo.
Use Cases
-
Carnegie Learning MATHia (K-12 mathematics, USA): the production deployment environment for BKT since the 1990s; the current platform serves over 700,000 students annually across 4,000+ US schools. Each of the platform’s approximately 500 mathematics knowledge components has BKT parameters estimated from millions of logged student–tutor interactions. The system’s mastery-based gating — holding students on a skill until P(L_t) ≥ 0.95 — directly operationalises Mastery Learning and is the basis for Carnegie Learning’s published learning gain evidence in algebra and geometry.
-
ASSISTments online homework platform: the canonical open-science deployment and benchmarking platform for knowledge tracing, with 170,000+ students across 800+ US schools. ASSISTments provides the reference datasets (ASSISTments 2009, 2015, 2017) used in all knowledge tracing benchmarking studies. BKT serves as the baseline against which Deep Knowledge Tracing, DKVMN, SAKT, AKT, and Transformer Architecture-based models are evaluated.
-
Duolingo (language learning, 500+ million users): employs a hybrid BKT-derived model for scheduling review exercises based on estimated word and grammar skill mastery, combining BKT’s probabilistic mastery tracking with Spaced Repetition scheduling algorithms (Half-Life Regression, Settles & Meeder 2016) to optimise long-term vocabulary and grammar retention. The A/B testing programme at Duolingo has demonstrated that BKT-informed scheduling reduces retention errors by 23% versus random scheduling baselines.
-
Computerised Adaptive Testing and Formative Assessment: BKT parameters inform item selection in formative assessment systems — when P(L_t) is near the mastery threshold, the system selects maximally diagnostic items to resolve uncertainty, mirroring the maximum-information item selection principle of Computerised Adaptive Testing grounded in Item Response Theory. The BKT and IRT approaches converge in high-stakes testing pipelines where online parameter updating is critical.
-
MOOC and corporate learning platforms: Coursera, edX, and enterprise learning platforms (Degreed, Cornerstone) use BKT-inspired mastery models to generate skill-mastery credentials, identify learners stuck on specific skills, and trigger instructor alerts when class-level BKT mastery distributions indicate systematic instructional gaps.
-
Special education and remediation: BKT’s per-skill parameter estimation makes it particularly effective for identifying anomalous learning patterns — very high slip rates suggesting test anxiety, very high guess rates indicating strategic guessing without comprehension, extremely low learn rates indicating prerequisite skill gaps — enabling targeted human intervention and Domain Model-informed remediation.
Academic Context
BKT emerged from the ACT-R cognitive modelling tradition at CMU. The foundational 1995 Corbett & Anderson paper in User Modeling and User-Adapted Interaction remains one of the most cited papers in educational computing, with over 3,500 citations as of 2026. The International Conference on Educational Data Mining (EDM, first held 2008) and the AIED conference have both treated BKT as the reference baseline for new knowledge tracing models. Key theoretical debates in the BKT literature include: identifiability and degeneracy of parameter estimates (Baker et al., 2008, 2010); the value of individualised vs. population-level parameter estimation (Yudelson et al., 2013); the relationship between BKT and Item Response Theory (Ruopp et al., arXiv:1803.05926, 2018); fairness and demographic bias in BKT-driven adaptive systems (van der Graaf et al., 2022; EDM 2025); and the relative merits of interpretable BKT versus Deep Knowledge Tracing for high-stakes deployment. The pyBKT library (Badrinath & Pardos, MDPI 2023) democratised BKT implementation, providing GPU-accelerated Python implementations of standard and extended BKT variants used in over 200 research groups.
Current Landscape (2026)
In 2026, BKT occupies a distinctive production niche: no longer state-of-the-art on held-out prediction benchmarks — Deep Knowledge Tracing and Transformer Architecture-based successors (SAKT, AKT, DIMKT) achieve 5–15% higher AUC by capturing cross-skill dependencies and richer temporal dynamics that BKT’s per-skill independence assumption ignores — but dominant in production ITS deployments due to interpretability, data efficiency, and regulatory compliance. The EDM 2025 conference explicitly highlighted a fairness track following multiple studies documenting demographic disparities in BKT-driven systems; the “Fairness of Bayesian Knowledge Tracing for Math Learners of Different Reading Ability” paper (EDM 2025) demonstrates that reading ability confounds BKT mastery estimates in mathematics, producing inequitable gating decisions. B²KT frameworks treating BKT parameters as per-student posteriors directly address this by naturally individualising instruction and producing more equitable curricula across reading ability and demographic groups. LLM integration has created the hybrid architecture most deployment teams favour for 2026: BKT manages skill-state estimation and mastery-based pacing (interpretable, auditable), while Large Language Model backends power Socratic dialogue, explanation generation, and dynamic problem creation — as in Carnegie Learning’s MathGPT and Khanmigo. Neural-Symbolic Knowledge Tracing (arXiv:2604.08263, 2026) injects Knowledge Graph prerequisite structure into deep Learner Models, offering a convergence path toward models that combine BKT’s structural grounding with DKT’s predictive power. The Language Bottleneck Model for Qualitative Knowledge State Modeling (arXiv:2506.16982, 2026) uses language models to produce qualitative knowledge state descriptions rather than binary mastery judgements — a potential paradigm shift toward semantically richer learner models that go beyond BKT’s binary latent state.
UK Context
-
University College London (UCL) Knowledge Lab: Dr Mutlu Cukurova and colleagues research multimodal learner modelling combining BKT-style mastery tracking with facial expression, speech, and physiological signals from Affective Computing, enabling affect-aware Adaptive Learning beyond response-only BKT. UCL is a partner in the EU-funded BOOST project evaluating adaptive learning and Knowledge Component Model-based mastery tracking at scale in European higher education, and leads UK policy-oriented research on whether adaptive systems serve or disadvantage learners with special educational needs or from low-income backgrounds — directly informing UK DfE evidence reviews.
-
Open University (Milton Keynes): the OU’s Learning Analytics team, working with Jisc, deploys BKT-inspired skill mastery tracking for the OpenLearn platform, combining it with Learning Analytics dashboards that surface skill-level mastery estimates to tutors and personal advisors. The OU’s scale — 170,000+ students on 600+ courses — provides unique datasets for large-scale BKT parameter estimation, fairness evaluation across demographic groups, and longitudinal validation of BKT-based mastery predictions against delayed assessment outcomes.
-
Century Tech (London): Century’s adaptive platform, deployed in 1,500+ UK schools, uses BKT-derived mastery signals combined with deep learning content recommendation to generate micro-lesson pathways at sub-skill granularity within mathematics, English, and science. The platform tracks mastery across hundreds of knowledge components per subject, with granularity distinguishing between specific misconception types within skills (e.g., sign-error vs. procedure-order errors in algebra). Century has published RCT evidence of 15–35% additional learning gain in UK secondary schools, constituting the strongest UK-based evidence for BKT-informed Adaptive Learning.
-
University of Edinburgh: Edinburgh’s AI in Education group within the School of Informatics works on Bayesian learner modelling and uncertainty-aware knowledge tracing for dialogue-based tutoring. Collaboration with Heriot-Watt University’s Interaction Lab applies BKT-style skill tracking to phonological and lexical knowledge components in second-language acquisition spoken dialogue tutors.
-
Regulatory context: the UK DfE’s Generative AI in Education guidance (2024) requires transparency about algorithmic content sequencing; GDPR Article 22 is particularly relevant where BKT mastery gating has consequential effects — blocking a student’s advance to the next topic without teacher override constitutes an automated decision subject to transparency, contestability, and Explainable AI requirements. Federated Learning approaches for cross-institutional BKT parameter estimation without centralising student data are being piloted by Jisc.
Future Directions (2026–2030)
-
Equity-by-design BKT: B²KT and successor frameworks that learn student-specific parameters online will supersede population-level parameter sets encoding demographic confounds; equity auditing toolkits for BKT-driven systems are under development by the EDM fairness working group, targeting deployment as standard practice by 2028.
-
BKT–LLM hybrid architectures: BKT provides the interpretable mastery-state backbone while Large Language Model backends generate diagnostic questions calibrated to P(L_t), adaptive explanations targeting specific misconceptions inferred from slip patterns, and natural-language transparency reports — closing the Explainable AI gap that prevents broader deployment in regulated UK and EU educational contexts.
-
Federated BKT for privacy-preserving cross-institutional deployment: fitting BKT parameters across distributed institutional datasets without centralising raw student response logs under Federated Learning protocols, enabling cross-institutional parameter sharing in compliance with GDPR and FERPA — piloted by Jisc’s federated analytics infrastructure.
-
Multi-skill BKT with Graph Neural Network structure: using the prerequisite graph of a curriculum as a structural prior over BKT parameter sharing — skills with similar prerequisite dependencies share parameters — enabling robust estimation for low-frequency skills with sparse interaction logs.
-
Affective Computing integration: adaptive tutoring platforms combining BKT skill mastery with physiological signals from wearable devices to detect frustration and cognitive overload, enabling affect-aware mastery gating that adjusts modality and pace as well as content.
-
Neuroimaging calibration: emerging research using fMRI and EEG during mathematics learning at CMU, UCL, and Stanford explores whether neural signatures of skill acquisition can calibrate or validate BKT P(T) estimates, providing neuroscientific grounding for the cognitive assumptions embedded in the model.
Relationship to Item Response Theory and Psychometrics
Bayesian Knowledge Tracing and Item Response Theory (IRT) are the two dominant probabilistic frameworks in Educational Technology, and their comparison illuminates fundamental choices in learner modelling. IRT models the probability of a correct response as a function of latent student ability θ and item parameters (difficulty b, discrimination a, guessing c), using a logistic function P(correct|θ,a,b,c) = c + (1-c)/(1+exp(-a(θ-b))). The latent variable θ is continuous (on a scale analogous to a z-score), stable within a testing session, and not assumed to change between items — making IRT a static measurement model designed for the snapshot assessment context. BKT, by contrast, models a binary latent state per skill that is explicitly designed to change between observations: its learning transition P(T) is the mechanism by which practice drives mastery acquisition, making BKT a dynamic learning model designed for the formative, mastery-based tutoring context.
The conceptual relationship between the two frameworks is illuminated by the Learning meets Assessment paper (Ruopp et al., arXiv:1803.05926, 2018): BKT’s guess and slip parameters P(G) and P(S) play roles analogous to IRT’s guessing parameter c and 1-c_slip respectively, while BKT’s per-skill binary mastery state corresponds to a discretised version of IRT’s continuous ability dimension restricted to a single skill. The key difference is that BKT models the learning process explicitly (P(T) captures how assessment opportunities drive state transitions), while IRT treats ability as fixed during the test and requires separate longitudinal models to capture learning over time. KT-IDEM (Pardos & Heffernan, 2011) is the most direct bridge, extending BKT with item-specific emission probabilities that mirror IRT’s item characteristic curve without abandoning the HMM temporal dynamics. More recent hybrid IRT-KT approaches (CAT-ML hybrids, 2024–2025) use IRT ability estimates as features in neural knowledge tracing models, combining IRT’s psychometric precision in ability estimation with DKT’s temporal dynamics modelling.
From a Psychometrics perspective, BKT can be interpreted as a special case of multistage testing or adaptive mastery testing where the stopping rule is probabilistic (advance when posterior mastery probability ≥ threshold) rather than fixed-length. Computerised Adaptive Testing theory (Wainer et al., 2000) provides theoretical tools — stopping rules, bias correction, conditional standard error of measurement — that are directly applicable to BKT’s mastery decision, and the cross-pollination between the CAT and ITS communities has been productive: several platforms now implement hybrid BKT-IRT systems that use IRT for item difficulty estimation and BKT for mastery state tracking, combining the strengths of both frameworks.
BKT in Production: Implementation Notes
The practical implementation of Bayesian Knowledge Tracing in production tutoring platforms involves a series of engineering and pedagogical decisions that the academic literature often glosses over but that significantly affect real-world performance. Parameter estimation frequency: in production systems, BKT parameters are typically estimated offline in periodic batch updates (weekly or monthly) from accumulated interaction logs, rather than in real time. Online re-estimation as new student data accumulates is theoretically attractive but computationally expensive at scale and can introduce instability when individual student populations shift. Skill granularity: the appropriate level of granularity for knowledge components significantly affects BKT behaviour — too coarse (treating all of algebra as one KC) and the model cannot identify specific gaps; too fine (treating every minor procedural variant as a separate KC) and insufficient data accumulates per KC for reliable parameter estimation. Carnegie Learning’s experience suggests that 50–500 KCs per course represents the practical sweet spot. Mastery threshold calibration: the conventional 0.95 threshold is often calibrated in practice to trading off advancement speed against mastery assurance — lower thresholds (0.85–0.90) advance students faster and reduce disengagement from repetitive practice but increase the probability of premature advancement and downstream difficulty. Multi-KC dependencies: standard BKT assumes independence across KCs; in practice, teachers and curriculum designers know that failure to master prerequisite skills will predictably slow mastery of dependent skills, and this signal — visible in the BKT parameter learning rates across the prerequisite graph — is often used to trigger prerequisite remediation rather than continued practice on the dependent skill.
The pyBKT library (Badrinath & Pardos, 2021; MDPI 2023) has standardised production-quality BKT implementation in Python, providing:
-
Vectorised parameter estimation for hundreds of KCs simultaneously using GPU-accelerated EM
-
Multiple BKT variants: standard, Individualised (per-student P(L₀)/P(T)), KT-IDEM (per-item parameters), BKT+Forget (forgetting parameter)
-
Cross-validation and model selection utilities for choosing between BKT variants
-
Mastery prediction and curriculum sequencing utilities that translate BKT posteriors into adaptive decisions
-
Integration with ASSISTments data format and the standardised KT benchmark evaluation pipeline
The library’s adoption across 200+ research groups and multiple commercial platforms has de-facto standardised BKT implementation, making pyBKT the reference implementation for reproduction and benchmarking — analogous to the role scikit-learn plays for general Machine Learning algorithms.
Research and Literature
- Corbett, A.T. & Anderson, J.R. (1995). Knowledge tracing: Modelling the acquisition of procedural knowledge. User Modeling and User-Adapted Interaction, 4, 253–278.
- Anderson, J.R., Corbett, A.T., Koedinger, K.R., & Pelletier, R. (1995). Cognitive tutors: Lessons learned. Journal of the Learning Sciences, 4(2), 167–207.
- Yudelson, M.V., Koedinger, K.R., & Gordon, G.J. (2013). Individualized Bayesian Knowledge Tracing Models. Proceedings of AIED 2013. Springer.
- Baker, R.S.J.d., Corbett, A.T., & Aleven, V. (2008). More accurate student modeling through contextual estimation of slip and guess probabilities in BKT. Proceedings of ITS 2008. Springer.
- Pardos, Z.A. & Heffernan, N.T. (2011). KT-IDEM: Introducing item difficulty to the knowledge tracing model. Proceedings of UMAP 2011. Springer.
- Piech, C. et al. (2015). Deep Knowledge Tracing. NeurIPS 2015.
- Zhang, J., Shi, X., King, I., & Yeung, D-Y. (2017). Dynamic Key-Value Memory Networks for Knowledge Tracing. WWW 2017.
- Pandey, S. & Karypis, G. (2019). A Self-Attentive model for Knowledge Tracing (SAKT). EDM 2019.
- Ghosh, A., Heffernan, N., & Lan, A.S. (2020). Context-Aware Attentive Knowledge Tracing (AKT). KDD 2020.
- Settles, B. & Meeder, B. (2016). A Trainable Spaced Repetition Model for Language Learning. ACL 2016.
- van der Graaf, J. et al. (2022). Bayesian-Bayesian Knowledge Tracing for more equitable tutoring. EDM 2022.
- Badrinath, A., Wang, F., & Pardos, Z. (2021). pyBKT: An Accessible Python Library of Bayesian Knowledge Tracing Models. EDM 2021.
- Badrinath, A. & Pardos, Z. (2023). An Introduction to Bayesian Knowledge Tracing with pyBKT. Multimodal Technologies and Interaction, 5(3), 50. MDPI.
- Van de Sande, B. (2013). Properties of the Bayesian Knowledge Tracing Model. Journal of Educational Data Mining, 5(2), 1–10.
- Bloom, B.S. (1984). The 2 sigma problem: The search for methods of group instruction as effective as one-to-one tutoring. Educational Researcher, 13(6), 4–16.
- Lord, F.M. (1980). Applications of Item Response Theory to Practical Testing Problems. Lawrence Erlbaum.
- Ruopp, M. et al. (2018). Learning meets assessment: On the relation between IRT and BKT. arXiv:1803.05926.
- Anonymous (2022). Equity and Fairness of Bayesian Knowledge Tracing. arXiv:2205.02333.
- Anonymous (2025). Fairness of BKT for Math Learners of Different Reading Ability. EDM 2025 Proceedings.
- Anonymous (2025). Sparse Binary Representation Learning for Knowledge Tracing. arXiv:2501.09893.
- Anonymous (2025). Deep Learning Based Knowledge Tracing: A Review of the Literature. BDAIE 2025. ACM.
- Anonymous (2025). Next Token Knowledge Tracing: Exploiting Pretrained LLM Representations to Decode Student Behaviour. arXiv:2511.02599.
- Anonymous (2025). Interpretable Knowledge Tracing via Transformer-Bayesian Hybrid Networks. Applied Sciences, 15(17), 9605. MDPI.
- Anonymous (2025). Can LLMs Generate Accurate Bayesian Networks to Enhance Knowledge Tracing? AIED 2025. Springer.
- Anonymous (2025). Knowledge Tracing Based on Learner Fatigue State. Complex & Intelligent Systems. doi:10.1007/s40747-025-01831-x.
- Anonymous (2026). Neural-Symbolic Knowledge Tracing: Injecting Educational Knowledge into Deep Learning. arXiv:2604.08263.
- Anonymous (2026). Language Bottleneck Models for Qualitative Knowledge State Modeling. arXiv:2506.16982.
- Luckin, R. et al. (2016). Intelligence Unleashed: An Argument for AI in Education. Pearson Education.
Historical Development and Intellectual Lineage
Bayesian Knowledge Tracing’s intellectual lineage traces through converging traditions in educational psychology, cognitive science, and mathematical statistics. The underlying pedagogical vision originates with Benjamin Bloom’s (1984) identification of the Two-Sigma Problem: that one-to-one human tutoring produces learning outcomes two full standard deviations above conventional classroom instruction, placing the average tutored student at the 98th percentile of conventionally-taught peers. The implication for computational Adaptive Learning was clear — if the mastery-checking and pacing intelligence of an expert tutor could be automated, it could be delivered to every learner simultaneously. This vision shaped the whole Cognitive Tutor programme at CMU.
The immediate precursor to BKT was the ACT* (Adaptive Control of Thought) theory developed by John Anderson in the 1970s–1980s, which proposed that procedural skills are acquired as production rules through a three-stage process: declarative knowledge (knowing that), compilation (converting declarative knowledge into procedural rules through practice), and tuning (strengthening and specialising rules through error-driven correction). This cognitive architecture provided the theoretical basis for the Cognitive Tutor’s component-level skill modelling: each production rule in the expert model corresponded to a skill that could be independently mastered, directly analogous to the BKT per-knowledge-component independence assumption.
The statistical innovation that Corbett and Anderson made in 1994–1995 was recognising that ACT*-inspired skill acquisition could be modelled as a Hidden Markov Model — a well-established probabilistic framework for systems with latent states generating observable outputs — and that Bayesian updating on the posterior over this latent mastery state provided a principled, online, computationally efficient mechanism for tracking the cognitive state of an individual learner from their response history. This was not the first use of Bayesian inference in education (Item Response Theory dates to Lord’s 1980 formalisation), but it was the first application of sequential Bayesian Inference with a dynamical learning model to the real-time tutoring context, producing a genuinely adaptive system that updated its beliefs after every single response rather than requiring a batch assessment.
The deployment of BKT in the CMU Algebra Cognitive Tutor beginning in the early 1990s — and the subsequent large-scale controlled evaluation by Koedinger and Anderson (1997) demonstrating statistically significant learning gains in a school-based study of 470 Pittsburgh students — provided the first empirical validation of BKT-informed Adaptive Learning at classroom scale. Carnegie Learning’s commercialisation of the Cognitive Tutor beginning in 1998, and the subsequent adoption by over 4,000 US schools, made BKT the most widely deployed Machine Learning model in K-12 education by an order of magnitude. The availability of the resulting large-scale response datasets — tens of millions of student–skill interactions logged with millisecond temporal resolution — then enabled the Educational Data Mining (EDM) research community, when it formalised in 2008, to use Carnegie Learning’s data for rigorous retrospective analysis that deepened understanding of BKT’s properties, limitations, and alternatives.
The publication of Deep Knowledge Tracing (Piech et al., NeurIPS 2015) marked the beginning of a new phase in which BKT became a baseline rather than the state of the art. DKT’s LSTM model, trained on the ASSISTments dataset, outperformed BKT by 6–25% AUC across four benchmark datasets by capturing cross-skill dependencies and long-range temporal effects that BKT’s per-skill independence assumption cannot represent. Subsequent models — DKVMN (Zhang et al., 2017), SAKT (Pandey & Karypis, 2019), AKT (Ghosh et al., 2020) — each advanced predictive performance through Attention Mechanisms, Graph Neural Networks, and multi-scale temporal representations. Yet none displaced BKT in production deployments, because the properties that make BKT attractive for deployment — interpretability, data efficiency, explicit mastery threshold, and alignment with Mastery Learning pedagogy — are orthogonal to the held-out AUC metrics on which neural models are evaluated. The research-deployment gap in knowledge tracing is one of the clearest examples in Educational Technology of the distinction between optimising for a benchmark metric and optimising for the goals of the deployed system.
LLM Integration and Current Transformation (2024–2026)
The emergence of Large Language Models as a platform for educational dialogue has created the most significant opportunity for BKT since its initial development: not as a replacement but as a complementary architecture in which BKT’s structured mastery tracking drives LLM-powered adaptive conversations. The resulting hybrid systems separate two previously conflated functions: knowledge state estimation (which BKT handles through principled Bayesian Inference on response sequences) and pedagogical dialogue generation (which LLMs handle through fluent natural-language conversation calibrated to the learner’s stated understanding). Carnegie Learning’s MathGPT, Khan Academy’s Khanmigo, and Duolingo Max each implement variants of this architecture.
The integration operates in two directions. Downstream from BKT to LLM: the current BKT posterior P(L_t) for each knowledge component is passed as context to the LLM, allowing the dialogue tutor to generate explanations targeted to the learner’s specific mastery state (“Based on your last five responses, you understand fraction multiplication but appear to struggle with equivalent fraction identification — let me show you a different way to visualise denominators”). This grounding prevents the LLM from hallucinating the learner’s knowledge state or providing explanations appropriate for a different level of mastery. Upstream from LLM to BKT: LLM-generated practice items, worked examples, and diagnostic questions can be tagged with skill labels from the Knowledge Component Model and fed into the BKT response-observation pipeline, expanding the effective item bank beyond pre-authored content and addressing the item depletion problem that limits classical Computerised Adaptive Testing at high practice volumes.
The language bottleneck approach (arXiv:2506.16982, 2026) goes further, using LLMs to generate qualitative knowledge state descriptions — “The learner appears to have procedural fluency with standard cancellation but may lack conceptual understanding of why the procedure is valid” — as a richer complement to BKT’s binary P(L_t) estimate. Such qualitative descriptions better support pedagogical decisions about when to provide worked examples vs. conceptual explanations vs. additional practice problems, and are more interpretable to learners than probability estimates. Neural-Symbolic Knowledge Tracing (arXiv:2604.08263, 2026) injects Knowledge Graph prerequisite structure as symbolic constraints into deep learning learner models, producing models that inherit BKT’s structural grounding (prerequisite constraints, skill independence where appropriate) while leveraging Transformer Architecture representation learning for rich temporal and inter-skill modelling.
The risks of LLM integration are real and pedagogically significant. LLMs trained on internet text may generate confident but incorrect mathematical explanations, historical claims, or reasoning steps, creating misinformation at precisely the moment when the learner’s knowledge state is most receptive — making LLM hallucination in educational contexts potentially more harmful than in general information retrieval. The 2025 Frontiers in AI Education study found that vanilla Large Language Model prompts used as tutors produced Socratic dialogue (guiding discovery) in only 31% of interactions vs. 78% for pedagogically fine-tuned variants — highlighting that raw LLM capability without pedagogical constraints tends toward delivering answers rather than scaffolding understanding. BKT’s explicit mastery criterion provides one safeguard: even when an LLM generates a persuasive but incorrect explanation, the subsequent response pattern under BKT will eventually reveal that understanding has not been achieved, triggering further instruction.
Benchmark Datasets
The knowledge tracing community has converged on a small set of canonical open datasets for benchmarking. The ASSISTments 2009 dataset (4,151 students, 110 skills, 525,534 interactions) from the ASSISTments platform at WPI remains the primary benchmark; standard BKT achieves AUC ≈ 0.69 on this dataset while Deep Knowledge Tracing achieves AUC ≈ 0.82 and attention-based models (AKT) achieve AUC ≈ 0.85. The KDD Cup 2010 Algebra dataset (575 skills, 8.9 million interactions from Carnegie Learning’s Cognitive Tutor) is the largest publicly available Knowledge Component Model-tagged dataset, used for individualised BKT evaluation. The Statics 2011 dataset from CMU’s Open Learning Initiative — a small but carefully curated dataset with 189 students and 1,224 questions across 85 skills in introductory statics — is used for evaluation of fine-grained BKT parameter estimation. ASSISTments 2015 (19,917 students, 100 skills, 708,631 interactions) provides a larger temporal split suitable for longitudinal validation. EdNet (Riiid, 2020) is the largest open Educational Technology dataset with 784,309 learners and 131,441,538 interactions in an English language preparation context, enabling large-scale evaluation of Neural Network-based knowledge tracing that would overfit on smaller corpora. Performance on these benchmarks: standard BKT ≈ 0.69 AUC; individualised BKT ≈ 0.72 AUC; BKT-LSTM ≈ 0.78 AUC; Deep Knowledge Tracing ≈ 0.82 AUC; SAKT ≈ 0.83 AUC; AKT ≈ 0.85 AUC; LLM-based knowledge tracing (2024-2026) reporting up to 0.88 AUC on ASSISTments by leveraging semantic understanding of question text beyond response patterns alone.
Comparative Analysis: BKT vs. Successor Models
The knowledge tracing landscape in 2026 comprises a spectrum of models spanning from BKT’s interpretable four-parameter Hidden Markov formulation to Large Language Model-based approaches that represent knowledge states implicitly through billions of parameters. Understanding where BKT sits in this spectrum — and when to prefer it over alternatives — requires comparing models across multiple dimensions: predictive performance, data efficiency, interpretability, computational cost, fairness, and alignment with pedagogical goals.
On held-out prediction benchmarks (AUC on ASSISTments 2009): standard BKT ≈ 0.69; individualised BKT ≈ 0.72; BKT-LSTM ≈ 0.78; Deep Knowledge Tracing LSTM ≈ 0.82; DKVMN ≈ 0.82; SAKT (self-Attention Mechanism) ≈ 0.83; AKT (context-aware attention) ≈ 0.85; LLM-based KT (2024–2026 papers) ≈ 0.86–0.88. This performance ordering is unambiguous: deeper models with more parameters and richer temporal representations consistently outperform BKT on the prediction task. However, the prediction task — given response history, predict the correctness of the next response — is a proxy for the actual deployment goal of producing mastery estimates that drive effective adaptive sequencing and pacing decisions. Evidence that better benchmark AUC translates to better adaptive tutoring outcomes is surprisingly sparse: the large-scale controlled studies that demonstrate learning gains (Carnegie Learning’s MATHia, DreamBox, Century Tech) overwhelmingly use BKT-based or BKT-derived mastery models, not neural knowledge tracing models, because the latter’s continuous prediction outputs do not map cleanly to the discrete mastery-gating decisions that operationalise Mastery Learning pacing.
On data efficiency: BKT converges to reliable parameter estimates with 20–50 observations per skill per student; individualised BKT requires 50–100 observations per student–skill pair for stable student-specific parameter estimation; Deep Knowledge Tracing requires thousands of observations to train reliably and is prone to overfitting on datasets with fewer than 10,000 total interactions. For low-frequency skills in a curriculum — particularly advanced or optional topics that few students reach — BKT is the only viable option, because DKT and its successors cannot be meaningfully trained on 15 observations. This data efficiency advantage is not a mere implementation convenience but a fundamental consequence of BKT’s parametric structure: four parameters per skill versus 175,000+ LSTM parameters in typical DKT implementations. The Transfer Learning problem — using knowledge about similar skills to initialise estimates for new or data-sparse skills — is more tractable for BKT (sharing P(T) across skills with similar cognitive complexity) than for DKT (requiring full network architecture sharing, which loses skill-specific representation quality).
On fairness and equity: as established by van der Graaf et al. (2022) and the EDM 2025 fairness track, both BKT and DKT encode demographic confounds from training data. BKT’s explicit parameter structure makes these confounds more diagnosable — if P(G) is systematically estimated higher for certain demographic groups, this is visible and correctable — while DKT’s implicit representations make demographic bias harder to detect and harder to remediate without adversarial debiasing techniques. B²KT’s per-student parameter treatment addresses BKT’s fairness limitation by making the demographic confounds sources of posterior variance rather than fixed parameter estimates, effectively individuating the model to each student’s characteristics.
On interpretability and governance: BKT is the uniquely interpretable option when deployed in regulated educational contexts requiring explanation, contestability, and teacher oversight. A BKT-based system can explain its mastery decision as “Your student has answered 12 questions on algebraic fractions; our model estimates 87% probability of mastery but is holding the student for 2 more practice opportunities due to 3 recent slip errors — click here to override” in a way that supports informed teacher judgment. DKT or Transformer Architecture-based models cannot provide this explanation without additional post-hoc interpretability layers that may not faithfully represent the model’s actual decision logic.
Empirical Evidence Base
The evidence base for BKT-informed Adaptive Learning effectiveness spans controlled laboratory studies, classroom randomised controlled trials, and large-scale platform data analysis. Bloom’s (1984) Two-Sigma Problem paper — the foundational empirical reference for the entire field — demonstrated the upper bound of what mastery-based individualised tutoring could achieve in controlled university psychology experiments, providing the aspirational target against which all adaptive tutoring systems are evaluated. Koedinger and Anderson’s (1997) 3-year longitudinal study of 470 Pittsburgh secondary school students using the Algebra Cognitive Tutor found statistically significant gains (effect size d = 0.4–1.0) vs. comparison conditions, providing the first classroom-scale validation of BKT-informed tutoring. A meta-analysis by Ma et al. (2014) across 107 controlled studies of Intelligent Tutoring Systems found a mean effect size of d = 0.66 vs. traditional instruction, consistent with Bloom’s predictions and attributable substantially to the mastery-tracking and pacing component that BKT provides.
Platform-scale evidence is more recent and methodologically more complex. Carnegie Learning published results from a large-scale quasi-experimental study (Ritter et al., 2007; VanLehn, 2011) showing that students using MATHia for the equivalent of one semester of algebra achieved 15–20% better post-test scores than comparison students. DreamBox Learning published a 2024 RCT with 12,000 K-5 students (n = 6,000 treatment, 6,000 control) demonstrating 0.26 SD gains in mathematics over one academic year attributable to the adaptive platform’s BKT-driven content selection. Century Tech (UK) published RCT evidence of 15–35% additional learning gain for UK secondary school students using the platform 3× per week for one term vs. matched controls, attributed primarily to the system’s granular Knowledge Component Model-based mastery tracking that identifies and targets specific misconceptions within skills rather than treating topics as monolithic units.
ASSISTments’ own platform experiments — using its A/B testing infrastructure across 170,000 students — have examined specific BKT model variants: Pardos & Heffernan (2011) demonstrated that KT-IDEM (item difficulty extension) improved BKT accuracy by 3–5% AUC on the platform; Wang et al. (2018) showed that adding response-time features to BKT observation models improved performance on fast/slow response patterns. These incremental improvements validate the BKT framework’s extensibility: the four-parameter HMM can be augmented with additional observations (response time, hint requests, error types) beyond binary correctness without abandoning the interpretable probabilistic structure.
Challenges and Open Problems
Despite BKT’s proven deployability, several unresolved challenges limit its effectiveness in real-world educational contexts. The knowledge component definition problem is arguably the most fundamental: BKT’s accuracy depends entirely on the quality of the Knowledge Component Model over which it operates, and there is no definitive algorithm for delineating knowledge components from curriculum content. Expert cognitive task analysis is expensive, subjective, and difficult to scale; data-driven KC discovery (Learning Factors Analysis, Koedinger et al. 2012) extracts KC models from response data but requires large datasets and may identify components that are statistically discriminable but pedagogically unmotivated. The forgetting problem in standard BKT — the assumption that mastery is permanent within a session — is violated in practice for skills that are rarely practised, where retention degays between sessions at rates that vary substantially across learners and skills. BKT+Forget extensions address this but add a fifth parameter that is hard to estimate reliably from sparse inter-session data, particularly when students practice skills at irregular intervals. The transfer and prerequisite problem — that mastering one skill should accelerate mastery of related skills — is structurally invisible to per-skill independent BKT models; addressing it requires either multi-skill BKT models with shared parameters or the Knowledge Graph-based extensions emerging in 2025–2026 research.
Ethical Considerations and Governance
The deployment of BKT-driven adaptive systems in educational contexts raises substantive ethical questions that have moved from academic concern to active regulatory engagement. The central tension is between the system’s need for granular interaction data — every response, hint request, response time, and navigation event — and the Data Privacy rights of learners, particularly minors. GDPR (UK and EU), FERPA (US), and COPPA (US, under-13s) impose strict requirements on educational data collection, storage, purpose limitation, and deletion; UK platforms must additionally comply with the Children’s Code (Age Appropriate Design Code) enforced by the ICO. The Adaptive Learning system’s inference that a learner has a specific knowledge gap — and the decision to hold them on that topic rather than advancing — constitutes algorithmic profiling of an educational kind; GDPR Article 22 provides a right not to be subject to solely automated decisions with significant effects, and the UK DfE’s 2024 guidance requires meaningful teacher oversight of mastery gating decisions.
Algorithmic bias is a particularly acute concern in BKT systems. The model learns P(G) and P(S) from historical data in which confounding demographic factors (socioeconomic status, reading ability in cross-curricular skills, English as additional language, neurodivergence) are entangled with measured performance. A student with high reading difficulty performing mathematics on a text-heavy platform will show elevated P(G) estimates and depressed P(T) estimates that reflect reading confounds rather than mathematical knowledge state — leading to inequitable gating decisions. EDM 2025’s dedicated fairness track and the Fairness of BKT paper (2025) directly address this: BKT models trained on heterogeneous populations may produce systematically inequitable decisions across demographic groups unless fairness constraints are applied during parameter estimation. B²KT’s per-student parameter treatment provides a partial solution by allowing reading-ability and other student-specific factors to be modelled as sources of parameter variation rather than noise, but requires sufficient per-student data to estimate student-specific posteriors reliably. Open Learner Model transparency — making P(L_t) and its basis visible and challengeable by learners — is recommended as an equity-promoting practice, allowing students to contest automated mastery decisions that may reflect confounded rather than genuine knowledge state estimates.
Key Terminology
-
Knowledge Component (KC): the unit of knowledge tracked by BKT; corresponds to a discrete procedural skill or declarative fact in the Knowledge Component Model, mapped from curriculum objectives via cognitive task analysis.
-
Mastery threshold: the P(L_t) value above which the system declares mastery and advances the learner; conventionally θ = 0.95 in Carnegie Learning deployments, reflecting a 95% posterior probability of having acquired the skill.
-
Slip rate P(S): probability of an incorrect response despite true mastery — captures test anxiety, careless errors, time pressure, and random lapses; typically estimated at 0.1–0.25 in well-calibrated BKT models.
-
Guess rate P(G): probability of a correct response without mastery — captures lucky guessing, multiple-choice format effects, and partial knowledge; should satisfy P(G) < 0.5 and P(G) + P(S) < 1 for a non-degenerate model.
-
Learn rate P(T): per-practice-opportunity probability of transitioning from unmastered to mastered state; high P(T) indicates a skill easily acquired within a session; low P(T) indicates slow accretion requiring many opportunities.
-
Prior knowledge P(L₀): probability of mastery before any interaction; encodes the expected fraction of learners who enter with prior knowledge of the skill from previous instruction or experience.
-
Degeneracy: pathological parameter configuration where P(G) + P(S) ≥ 1, making the model’s mastery predictions uninformative; Baker et al. (2010) showed approximately 20% of skills in unconstrained BKT fitting on real datasets exhibit degeneracy, requiring post-hoc correction.
-
Baum-Welch algorithm: the Expectation-Maximisation algorithm applied to Hidden Markov Models; alternates between computing expected state occupancies (E-step) and re-estimating parameters to maximise expected log-likelihood (M-step); the standard fitting procedure for BKT.
-
Open Learner Model (OLM): a transparency technique displaying the system’s current belief P(L_t) to the learner, enabling self-regulated learning and providing contestability of automated mastery decisions; recommended under UK DfE 2024 governance guidelines for AI in education.