← WORKING PAPERS
AOHA WORKING PAPER / 001 / v0.2SEPTEMBER 2026

From Answer Production to Thought Formation

Cognitive Responsibility and Dual Transfer in Generative-AI-Supported Education

AUTHORShinichi Asai AFFILIATIONHuman & Technology Management Laboratory, AOHA LAB TYPETargeted narrative synthesis and provisional conceptual framework

Abstract

Generative AI can improve performance while assistance is available, but its effects on later learning vary with the task, learner, system, and outcome measure. This targeted narrative synthesis examines selected experiments on generative-AI tutoring alongside foundational research on learning–performance distinctions, self-explanation, instructional assistance, transfer, and cognitive offloading. It argues that access alone is an inadequate unit of analysis. The more precise question is how cognitive responsibility is allocated among learner, AI, and teacher over time. We propose a provisional Cognitive Responsibility and Dual-Transfer model: define the cognitive operation to be learned; distinguish it from operations that may be offloaded; calibrate support to prior knowledge and task complexity; require reconstruction and verification; fade assistance; and assess both unaided capability transfer and responsible AI-augmented transfer. The model integrates established mechanisms rather than claiming that each component is novel. Its incremental contribution is to connect responsibility allocation with two forms of transfer and make the resulting claims testable. The framework has not been empirically validated as a unified intervention.

Keywords: generative AI, education, cognitive responsibility, transfer, metacognition, scaffolding, cognitive offloading, assessment

Review scope and method

This article is a targeted narrative synthesis, not a systematic review. Literature was identified in September 2026 through structured scholarly web searches, Crossref and publisher records, and backward citation chaining. Search concepts combined terms including generative AI, learning, transfer, AI tutor, self-explanation, worked examples, cognitive offloading, and assessment.

Selection prioritized directly relevant human-subject experiments, recent reviews, foundational learning theory, and policy guidance. Peer-reviewed evidence was preferred; a working paper was retained where it directly examined human tutors using AI and is identified as such. Opinion pieces and claims without inspectable research support were excluded. The search was not exhaustive, no formal risk-of-bias assessment was performed, and relevant studies may have been omitted. The synthesis therefore supports hypothesis generation rather than field-wide causal conclusions.

Better output is not the same as more learning.

Research on skill acquisition has long distinguished performance during practice from durable learning revealed after delay or under changed conditions.1 Generative AI makes this distinction unusually visible because the quality of a jointly produced artifact may increase even when the learner’s independent capability does not.

CONSTRUCTWORKING DEFINITIONILLUSTRATIVE MEASURE
Assisted performanceQuality achieved while AI support is available.Accuracy, code quality, or argument quality during use.
Unaided retentionKnowledge or skill available after support is removed.Delayed recall or reconstruction without AI.
Capability transferUse of the learned principle in a structurally different task with reduced or no AI.Near- and far-transfer tasks across content or context.2
Agency transferAbility to decide when, why, and how to use, constrain, verify, or reject AI in a novel task.Detection of flawed output and justified tool-use decisions.
Target cognitionThe mental operation the activity intends the learner to acquire.Problem representation, inference, explanation, or epistemic judgment.
Support cognitionAn operation that may be offloaded without defeating the learning objective.Formatting or routine calculation when these are not the target skill.
The relevant design question is not simply whether AI is present, but which cognitive responsibility moves to it, when, and with what consequence.

Selected evidence indicates conditional—not uniform—effects.

STUDY / CONTEXTDESIGN AND RESULTINTERPRETIVE LIMIT
Bastani et al.
Nearly 1,000 secondary students; mathematics; Türkiye
Cluster-randomized field experiment across four practice sessions. A GPT-4-based general interface raised assisted practice performance but was followed by 17% lower performance than control on an unassisted examination. A guardrailed tutor raised assisted performance without a detected examination penalty.3One school, one subject, and a short intervention. Removing detected harm is not evidence of superior durable learning.
Kestin et al.
N=194 eligible students; undergraduate physics; United States
Randomized crossover comparison of a purpose-built AI tutor and in-class active learning. The AI condition produced higher immediate post-test gains in less time and higher reported engagement.4Two lessons in one Harvard course; short-horizon outcomes. It does not establish general superiority to teachers or classrooms.
Wang et al.
900 tutors and 1,800 K–12 students; mathematics; United States
Preregistered trial of Tutor CoPilot, an AI support tool for human tutors. Access increased proximal topic mastery by four percentage points overall and by nine points for students of lower-rated tutors.5The intervention augmented human tutors, not direct student–AI learning; the main outcome was a proximal exit ticket.
Bassner et al.
N=275; introductory programming; Germany
Three-arm randomized trial comparing a scaffolded tutor, unrestricted ChatGPT, and no AI. Both AI groups improved exercise scores and reduced frustration, but neither improved conceptual learning or code comprehension.6A 90-minute task in one programming course; no delayed transfer test.
Deng et al.
69 experimental studies
Meta-analysis reported positive average effects on measured performance and affective-motivational outcomes.7Interventions and measures were heterogeneous; many outcomes did not isolate later unaided learning.

Finding: the reviewed studies include benefits, null effects, and a performance–learning dissociation. Synthesis: estimated effects belong to an intervention bundle—model, interface, prompts, sequencing, curriculum, teacher role, learner characteristics, and assessment—not to “AI access” in the abstract. Proposal: studies and schools should report assisted performance separately from later capability and agency transfer.

Useful difficulty must be calibrated, not romanticized.

Self-explanation can support conceptual learning and transfer,8 and constructive activity can produce more learning than passive reception.9 Yet requiring every learner to solve a full problem before receiving help would ignore the assistance dilemma.10 Novices facing high-complexity material may learn efficiently from worked examples,11 whereas prior generation can help under conditions that make failure productive rather than merely frustrating.12

The appropriate initial contribution may therefore be a prediction, subgoal, partial representation, confidence judgment, or explanation of one worked step. Support should become more explicit after diagnostic evidence of a prerequisite gap, then fade after the learner reconstructs the repaired reasoning. Cognitive offloading is likewise neither inherently harmful nor beneficial: it changes the operations performed internally and should be evaluated against the learning target.13

Relation to existing frameworks

EFFORT-AI already proposes an effort-preserving six-phase architecture—elicit, formulate, feedback, organize, reflect, and transfer—with load-adaptive support and accountability.14 The present paper does not claim to rediscover that sequence. Its narrower contribution is an allocation lens that can be applied within EFFORT-AI or other pedagogies.

FOUNDATIONWHAT IT CONTRIBUTESROLE HERE
Self-explanation / ICAPMechanisms of generative and interactive engagement.Explains why learner reconstruction may matter.
Assistance dilemma / worked examples / productive failureBoundary conditions for how much help to give and when.Prevents a universal “attempt first” rule.
EFFORT-AIA staged effort-preserving instructional architecture.Provides the closest prior framework and a comparison point.
Current proposalAllocation of target cognition plus capability and agency transfer.Makes responsibility and two transfer outcomes explicit across designs.

Cognitive Responsibility and Dual Transfer

  1. 01
    Specify target cognition

    Name the exact operation the learner should acquire: representation, inference, explanation, verification, judgment, or another defined capability.

  2. 02
    Allocate responsibility

    State what remains learner-owned, what AI may support or execute, and what requires teacher judgment.

  3. 03
    Calibrate assistance

    Use prior knowledge, task complexity, confidence, and observed errors to choose a micro-attempt, hint, worked fragment, or explicit explanation.

  4. 04
    Reconstruct and verify

    Ask the learner to rebuild the reasoning, diagnose the original error, and verify consequential claims with appropriate sources or ground truth.

  5. 05
    Fade support

    Reduce directiveness once the learner shows a usable representation and can explain the repaired step.

  6. 06
    Test dual transfer

    Measure both independent application and the capacity to use or challenge AI responsibly in a novel, ill-structured task.

A prompt history is not proof of cognition or authorship. Process data should be minimized and used formatively, then triangulated with explanation, delayed performance, and transfer. Learners should have clear retention rules and, where feasible, a non-AI route.

Operational case: proving π > 3.05

This illustration is a proposed instructional pathway, not evidence that the model works. The learning target is to connect the definition of π with geometric bounding and a justified inequality.

  1. 01
    Diagnostic contribution

    The learner defines π and predicts how an inscribed regular polygon could produce a lower bound. A novice may label a diagram or explain one worked step rather than attempt a full proof.

  2. 02
    Adaptive hint ladder

    The AI moves from a question about perimeter, to a diagrammatic cue, to a partially worked inequality only when the learner’s response shows the need.

  3. 03
    Reconstruction

    The learner restates why the polygon’s perimeter is less than the circumference and identifies which inequality establishes the bound.

  4. 04
    Verification

    Each geometric and numerical step is checked; the AI is also given one deliberately flawed step for the learner to diagnose.

  5. 05
    Capability transfer

    After delay and without AI, the learner proves a different lower bound or adapts the method using another polygon.

  6. 06
    Agency transfer

    With AI available, the learner compares two proposed proof strategies, rejects an invalid one, and justifies how the tool was used.

Four falsifiable hypotheses

H1 — Assistance timing. With instructional content and time held constant, contingent assistance following a bounded learner-generated response is hypothesized to produce higher delayed unaided retention and transfer than answer-first assistance, even if immediate assisted performance is equal or lower.

H2 — Expertise × complexity. Prior knowledge and task complexity are hypothesized to moderate the effect: novices facing high-complexity tasks should benefit more from a micro-attempt followed by worked fragments, while more knowledgeable learners should benefit more from sparse, progressively faded hints.

H3 — Reconstruction mechanism. Randomly requiring learners to reconstruct reasoning and diagnose an initial error after AI feedback is hypothesized to improve delayed transfer relative to feedback alone; reconstruction quality is a proposed mediator.

H4 — Dual transfer and calibration. Verification prompts plus faded support are hypothesized to improve both delayed unaided transfer and novel AI-augmented performance, including discrimination between correct and deliberately flawed AI responses.

A credible first test would preregister a factorial experiment crossing answer-first versus contingent assistance with reconstruction present versus absent, stratify by prior knowledge, fix model version and prompts, and measure immediate performance, delayed transfer, AI-augmented transfer, confidence calibration, cognitive load, motivation, and time.

Assessment must not become surveillance.

Restricting AI remains defensible in high-stakes assessment, privacy-sensitive work, some developmental contexts, and tasks where the tool would perform the precise operation being assessed. Elsewhere, responsible use requires attention to unequal access, language and cultural bias, hallucinated sources, age-appropriate consent, vendor retention, accessibility, teacher workload, and the availability of an equivalent non-AI route. UNESCO similarly frames educational AI as a human-centred design and governance problem.15

Reasoning traces should be concise, purpose-limited, and retained only as long as necessary. They are candidate indicators, not direct windows into thought. Oral explanation, multimodal response, and delayed transfer may provide fairer evidence for learners whose language, disability, or interaction style makes verbose prompting a poor proxy for understanding.

Limitations and conclusion

This selective review may omit relevant evidence and includes studies that differ in subject, age, geography, duration, AI role, and outcome horizon. The field changes rapidly, some results depend on specific model versions, and one central human-tutor study remains a working paper. The proposed model has not been tested as a unified intervention. Its terms and measures may require substantial adaptation for early-primary, multilingual, low-resource, accessibility, creative, and workplace-learning settings.

The resulting claim is deliberately conditional: generative AI may support durable learning when an instructional system makes cognitive responsibility explicit, calibrates assistance rather than maximizing it, returns consequential reasoning to the learner, and evaluates both unaided capability and responsible AI-augmented agency. Whether that claim holds is an empirical question.

Selected literature

  1. Soderstrom, N. C., & Bjork, R. A. (2015). Learning versus performance: An integrative review. Perspectives on Psychological Science, 10, 176–199.
  2. Barnett, S. M., & Ceci, S. J. (2002). When and where do we apply what we learn? A taxonomy for far transfer. Psychological Bulletin, 128, 612–637.
  3. Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, O., & Mariman, R. (2025). Generative AI without guardrails can harm learning: Evidence from high school mathematics. Proceedings of the National Academy of Sciences, 122, e2422633122.
  4. Kestin, G., Miller, K., Klales, A., Milbourne, T., & Ponti, G. (2025). AI tutoring outperforms in-class active learning: An RCT introducing a novel research-based design in an authentic educational setting. Scientific Reports, 15, 17458.
  5. Wang, R. E., Ribeiro, A. T., Robinson, C. D., Loeb, S., & Demszky, D. (2024). Tutor CoPilot: A human–AI approach for scaling real-time expertise. EdWorkingPaper 24-1054 / arXiv:2410.03017.
  6. Bassner, P., Lenk-Ostendorf, B., Beinstingel, R., Wasner, T., & Krusche, S. (2026). Less stress, better scores, same learning: The dissociation of performance and learning in AI-supported programming education. Computers and Education: Artificial Intelligence, 10, 100537.
  7. Deng, R., Jiang, M., Yu, X., Lu, Y., & Liu, S. (2025). Does ChatGPT enhance student learning? A systematic review and meta-analysis of experimental studies. Computers & Education, 227, 105224.
  8. Chi, M. T. H., Bassok, M., Lewis, M. W., Reimann, P., & Glaser, R. (1989). Self-explanations: How students study and use examples in learning to solve problems. Cognitive Science, 13, 145–182.
  9. Chi, M. T. H., & Wylie, R. (2014). The ICAP framework: Linking cognitive engagement to active learning outcomes. Educational Psychologist, 49, 219–243.
  10. Koedinger, K. R., & Aleven, V. (2007). Exploring the assistance dilemma in experiments with cognitive tutors. Educational Psychology Review, 19, 239–264.
  11. Atkinson, R. K., Derry, S. J., Renkl, A., & Wortham, D. (2000). Learning from examples: Instructional principles from the worked examples research. Review of Educational Research, 70, 181–214.
  12. Kapur, M. (2008). Productive failure. Cognition and Instruction, 26, 379–424.
  13. Risko, E. F., & Gilbert, S. J. (2016). Cognitive offloading. Trends in Cognitive Sciences, 20, 676–688.
  14. Liubertaitė, A. D. (2026). EFFORT-AI: An effort-preserving module for cognitively aligned human–AI learning. Frontiers in Education, 11, 1849821.
  15. Miao, F., & Holmes, W. (2023). Guidance for generative AI in education and research. UNESCO.

Research integrity and disclosure

Status: AOHA Working Paper v0.2. Not externally peer reviewed. No original human-participant experiment is reported. Funding: No funding declaration has been made for this version. Potential conflict: The author is the founder of AOHA LAB, which has a commercial interest in AI-related research, development, and education services. Data and code: Not applicable. Ethics approval: Not applicable to this literature-based paper.

AI-use disclosure: The concept and direction were provided by Shinichi Asai. Literature discovery, drafting, and editing were conducted with OpenAI Codex under the author’s direction. The manuscript underwent an AI-assisted internal critical review using three separately instructed reviewer perspectives covering learning science, human–AI interaction, and cross-disciplinary methodology and academic editing. The author remains responsible for final verification, interpretation, and publication. This process does not constitute independent human peer review.

Version 0.2 / 12 September 2026. Hypothesis-generating working paper; not externally peer reviewed and not a systematic review.