After reading this, you will understand how to identify behavioral signatures of learning in large-scale AI interaction logs. You will also see how assistant "scaffolding" influences user engagement. However, keep in mind that these observations are correlational. They describe how people behave, not necessarily how much they actually learn or retain.
As Large Language Models (LLMs) become deeply embedded in our daily cognitive work, a central tension has emerged. There is a growing concern regarding "cognitive offloading" (delegating mental effort to machines). This process might erode our own abilities. While much research focuses on designing AI specifically for education, we know very little about informal learning. This is learning that happens incidentally during everyday tasks.
A new study from researchers at the Hong Kong University of Science and Technology, Microsoft Research Asia, and Johns Hopkins University investigates this gap. They ask if everyday human–LLM interaction is purely about getting fast answers. Or, does it contain hidden opportunities for growth? By analyzing over 128,000 naturalistic conversations, the authors report that users do engage in learning-oriented behaviors. However, these moments of deep sense-making are selective rather than routine.
Mapping the signatures of informal learning
The researchers developed a method to translate theoretical constructs from learning science into measurable, turn-level behavioral signatures. Here, a "turn" is the basic unit of analysis, representing a single message from either the user or the assistant. This method produces a categorized map of cognitive effort and assistant support.
As shown in, the framework operates on two main axes.
On the user side, the method classifies turns into a hierarchy of cognitive engagement. These include passive receipt (simply acknowledging an answer), active use (applying or following up on information), and constructive engagement (the deepest level). Constructive engagement occurs when users elaborate, test, or revise ideas.
On the assistant side, the method distinguishes between "reference responses" and "scaffolded support." A reference response simply delivers the requested content or a standalone answer. In contrast, scaffolded support (assistance that structures or redirects the user's process) provides help like hints or explanations. This scaffolding aims to leave room for the user's own reasoning.
What You Need To Run It
To replicate this analysis, you need access to large-scale, naturalistic conversational datasets. The study utilizes three primary public corpora: WildChat, LMSYS Chat, and ShareChat. The scale is significant. It involves 128,569 conversations and nearly one million total turns.
The authors employed an LLM-assisted annotation pipeline to process these logs. This involves using high-capability models to perform semantic task filtering. They also use these models to assign turn-level labels. Because labeling nuanced human behavior is difficult, the authors implemented a rigorous validation workflow. This included both human-human agreement checks and human-LLM audits to ensure reliability. Code for the statistical analyses and the annotation workflow is reportedly available; see the paper for the canonical link at https://github.com/CinderD/Informal-learning-in-everyday-human-LLM-interaction.
How It Works
The core methodology relies on translating abstract learning theories into discrete, observable actions. Instead of trying to measure "knowledge gain," the authors focus on "process indicators."
The researchers categorize user engagement into three tiers. Passive engagement is akin to a passenger in a car simply watching the scenery. Active engagement is like a driver following GPS instructions. Constructive engagement is like a navigator questioning the route and suggesting alternatives. The study finds that while cognitive engagement appeared in 31.9% of user turns, the highest tier—constructive engagement—appeared in only 4.9% of turns [Table 1]. This means deep sense-making is a relatively rare occurrence in typical usage.
The analysis also examines "scaffolding." This is a technique used in teaching to provide temporary support. This support is gradually removed as the learner gains competence. The authors classify assistant support into various forms. These include feedback (M1), hinting (M2), and explaining (M4). By using integrated logistic regressions, the researchers modeled how these support forms interact with the user's immediate preceding state .
This reveals how the timing and style of help change user behavior.
How To Tell If It Worked
You will know the method is successfully identifying learning signatures when you observe specific patterns. The authors report that constructive engagement is "ecologically organized." Specifically, you should see higher rates of constructive engagement in coding-oriented tasks compared to writing tasks [Figure 2b]. This is because coding often forces users to deal with explicit errors and constraints.
Furthermore, successful identification of scaffolding should show a clear correlation with deeper participation. Conversations containing scaffolded support consistently show higher ratios of constructive user turns .
They also show more sustained interaction, meaning more turns follow the initial answer. Finally, at the granular level, you should see that certain support forms are more effective. For example, feedback and explanation are significantly more likely to be followed by a constructive user turn than direct instruction [Figure 5d].
Gotchas
There are several critical nuances to keep in mind. First, the study is observational. The authors are observing existing patterns in public logs. They are not assigning support in a controlled experiment. Therefore, you cannot claim that a specific type of AI response causes a user to learn more. You can only report that certain types of support are associated with more constructive behavior.
Second, the "visibility" of learning is a limitation. The authors note that constructive engagement is an observable proxy for learning. It is not a direct measure of it. A user might perform "constructive" actions without actually retaining the underlying concept long-term.
Third, the effectiveness of support is highly context-dependent. For example, the authors found that "explaining" (M4) consistently predicts constructive follow-up. This holds true regardless of the user's prior state. However, "feedback" (M1) was most effective specifically after a user had been in a passive state [Figure 5d]. A "one-size-fits-all" scaffolding strategy will likely fail to capture these subtle dynamics.
When This Is The Wrong Tool
This method is not intended for measuring actual skill acquisition or long-term memory retention. If your goal is to prove that an AI model increases test scores, this approach is insufficient. It is also not suitable for highly structured, curriculum-based environments. In those settings, learning objectives are predefined and the interaction is tightly controlled. This framework is designed specifically for the "wild," unstructured, and incidental learning that occurs during everyday problem-solving.
Figures from the paper
How this was made
Model: nvidia/Gemma-4-26B-A4B-NVFP4
Persona: academic_accessible
Template: tutorial
Refinement: 0
Pipeline: forge-1.1
Evaluator: nvidia/Gemma-4-26B-A4B-NVFP4
Score: 94% (passed)
Claims verified: 18 / 18
Model: nvidia/Gemma-4-26B-A4B-NVFP4
NVIDIA GB10 · 128 GB unified · NVFP4 · 100% local · $0 cloud
Tokens: 143,387
Wall-time: 326.2s
Tokens/s: 439.6