You can finish the worksheet, hand in the essay and spend an entire evening studying without being able to explain what you learned the next morning.
That should bother us more than it does.
We have made the visible parts of studying remarkably easy to count. Attendance. Assignments submitted. Pages covered. Hours at a desk. We can fill a dashboard with those numbers. The harder question is what happens when somebody closes the book and has to use the idea themselves.
Did they understand it? Can they explain why it works? Can they recognize when it does not apply? Will they remember enough to build on it next week?
That gap is why we are building plue. I want the work a learner puts in to lead somewhere they can actually feel: a problem that used to be confusing starts making sense, and eventually they can handle it themselves.
AI makes this more urgent. It can produce a convincing explanation, a finished assignment or a page of apparently excellent notes in seconds. A system that mistakes those outputs for understanding now has a much bigger problem.
There are three failures worth taking seriously: unequal access to the help that makes learning possible, weak connections between instruction and usable skills, and assessments that can reward work the learner could not have produced independently. The research puts numbers behind all three.
Start with the scale. UNESCO estimates that 273 million children and young people were out of school in 2024. The out-of-school rate was 36% in low-income countries and 3% in high-income countries. On-time upper-secondary completion reached 61% globally, but only 20% in low-income countries. These are different starting conditions for learning before anyone chooses a study technique. UNESCO’s 2026 education monitoring report.
Inside education systems, access to meaningful learning remains unequal. In PISA 2025, 48% of disadvantaged students were below Level 2 in mathematics, compared with 18% of advantaged students, on average across OECD countries. Disadvantage here means the bottom national socioeconomic quarter; advantage means the top quarter. This is a 30-percentage-point gap among assessed 15-year-olds, not a worldwide estimate covering every young person. OECD, PISA 2025 results.
Family resources, health and wider social conditions contribute to these differences. Schools cannot control all of them. But a learner's need for another explanation does not become smaller because their family cannot buy extra help.
Think about what useful support actually involves. Somebody notices the precise step you missed. They ask you to try it. Your answer reveals a misconception. They change the explanation and check again. That takes attention, and attention is unevenly distributed.
The difficulty is structural. UNESCO projects that 44 million additional primary and secondary teachers are needed by 2030 to achieve universal education at those levels. This is a staffing requirement toward a future goal, not a count of current vacancies. Asking existing teachers to give every learner unlimited individual attention does not make that attention available. UNESCO Global Report on Teachers.
An app will not fix poverty, school exclusion or teacher shortages. The part we can work on is specific: making useful help easier to reach when you are trying to understand something, and making that help responsive to what you do next.
The second failure is what happens between receiving instruction and developing a capability you can use.
The OECD's 2022–23 adult-skills assessment found that 26% of adults had low literacy proficiency, 25% had low numeracy proficiency and 29% had low adaptive problem-solving proficiency, across participating OECD countries and economies. These figures refer to adults aged 16–65 at Level 1 or below. They do not mean complete illiteracy, and adult skills reflect much more than schooling alone. They do show how much room there is between having participated in education and having strong foundational capabilities. OECD Survey of Adult Skills.
Instruction can make a substantial difference. A meta-analysis of 225 studies in undergraduate STEM found better assessment performance under active learning than traditional lecturing. Across studies reporting failure rates, the averages were 21.8% with active learning and 33.8% with traditional lectures. That is a 12-percentage-point difference, although the studies varied in design and setting. Freeman and colleagues, PNAS.
A separate published meta-analysis of 89 randomized tutoring experiments found an average improvement of 0.288 standard deviations. A standard deviation describes a score difference relative to the spread of scores in the study; it is not a percentage gain. This evidence concerns human tutoring programs, with different designs and results. Nickow, Oreopoulos and Quan, American Educational Research Journal.
These findings give us a serious reason to care about what the learner does during instruction. Listening to an explanation, attempting a problem and receiving feedback are different activities. We should be careful about treating them as interchangeable minutes of studying.
The same care belongs in conversations about qualifications. In an OECD analysis using 2023 data, 36% of employed adults aged 25–65 reported having skills above or below what their jobs required, across participating economies and excluding self-employment. That includes both over-skilling and under-skilling. It does not mean 36% were incapable of doing their jobs. Hiring practices, available jobs and working conditions contribute to the mismatch alongside education. OECD, A Skills-First Labour Market.
A qualification still has real value. Across OECD countries, tertiary-educated adults aged 25–64 working full-time and full-year earned 54% more on average than those with upper-secondary education, using 2023 or the latest available country data. That is an observed earnings difference, not a guaranteed personal return after costs. Anyone claiming education has become worthless needs to account for evidence like this. OECD, Education at a Glance 2025.
My concern is the promise we attach to the process. Complete these steps, collect the credential, and assume the understanding followed. Learners deserve a clearer view of what they can actually do and where they still need help.
Then AI arrives and makes the third failure visible: the quality of the submitted answer can separate from the capability of the person submitting it.
At the University of Reading, researchers inserted 63 entirely GPT-4-written submissions into five psychology modules alongside real student work. 59 of the 63, or 93.65%, were not flagged for any academic-practice concern. AI submissions achieved higher median grades in four of the five modules. This was one department using at-home assessments, not a global measure of cheating. It demonstrated that normal marking could reward fabricated student performance. Scarfe and colleagues, PLOS ONE.
The problem goes deeper than detection. In a randomized study at one Turkish high school, students using an unrestricted GPT-4 interface performed 48% better during assisted practice than the control group. On the subsequent exam without AI, they performed 17% worse, equivalent to 5.4 percentage points on the exam's 0–100 scale. A guided tutor produced larger assisted-practice gains and avoided that unaided-exam penalty, but did not establish an unaided-exam improvement. Bastani and colleagues, PNAS.
This was a short experiment in one school. It cannot tell us what every learner will experience over years. It can tell us something uncomfortable about measurement: a tool can help someone look more successful while leaving them less prepared to perform alone.
That is the failure I most want us to avoid at plue. A learner should not finish a session with an impressive answer and a false sense that the understanding belongs to them.
AI can also help when instruction is designed carefully. In a randomized crossover study involving 194 eligible Harvard physics undergraduates, a bespoke AI tutor produced a median immediate post-test score of 4.5 out of 6, compared with 3.5 out of 6 after classroom active learning. The experiment covered two lessons, so it does not establish long-term retention or that any chatbot will deliver the same result. Kestin and colleagues, Scientific Reports.
In nine public schools in Benin City, Nigeria, a six-week teacher-supervised, after-school AI program improved performance on a regular curricular English exam by 0.206 standard deviations. There was substantial attrition, and the comparison did not provide equivalent extra tutoring time without AI. The benefit belongs to the combined program, including its teachers and additional instruction. World Bank randomized evaluation.
So the design matters. What does the learner attempt? When does help appear? What gets checked afterward? Those questions are more useful than treating access to AI as an education strategy on its own.
In the desktop version of plue, we start with the material in front of you. With the permissions you choose, plue can use selected text and supported document context to begin a lesson where you got stuck. Relevant context from what you have been working on can help connect the explanation to the task. Click the orb when you want to ask a question out loud.
That interaction matters because confusion is usually specific. You might understand most of a page and get stuck on one relationship, one term or one step in a calculation. Starting there gives the explanation a purpose.
The lesson can use a visual instrument you manipulate. You can change a value and watch a plot respond, move a point in a geometric construction or explore a transformation. The picture gives you something to reason about while the explanation develops.
Then you have to make an attempt. A check might ask you to point to the steepest part of a curve, set a value, predict an outcome or write an expression. For supported answer types, plue checks the response and gives feedback. The point is to make your thinking visible enough that help can respond to it.
Your answers also affect what comes back for review. The desktop app records lesson progress and answer history, then uses an answer-based scheduling model to suggest when to revisit an idea. Reading a lesson and answering its question are recorded differently. Estimated recall and review dates are useful planning signals, with uncertainty attached; they are not guarantees about your memory.
These are current product behaviors. The larger learning loop we are building connects them more closely to the learner's course, goals and evidence over time.
We call the course context the Course Brain: the curriculum, supplied material, assessments, deadlines and sources that help establish what the learner is working toward. Coverage depends on the material and supported course context; we cannot claim to know every curriculum automatically.
The Director chooses a next teaching move, such as checking a prerequisite, explaining a step, offering a hint or asking for practice. The Tutor delivers that interaction. The Evaluator records what a supported learner response tells us, so the following step can take that evidence into account.
We are still connecting these parts. The complete adaptive course system is not fully deployed, and evaluating more kinds of learner responses reliably is part of the work ahead.
The studies in this article are also not trials of plue. They inform the questions we ask and the behaviors we prioritize. They do not establish that our product delivers their results. We need our own evidence about whether learners retain understanding, apply it to unfamiliar problems and become more capable without assistance.
There are limits beyond product design, too. The ITU estimated that 2.2 billion people remained offline in 2025. Internet use reached 94% in high-income countries and 23% in low-income countries. Any claim that everyone can simply replace education with an online tutor ignores who can reach that tutor in the first place. ITU Facts and Figures 2025.
We should keep the ambition large and the claims specific. Learning for a course, for work or for curiosity should give you more control over what you can understand and do. Getting an answer faster is valuable when it supports that. It becomes a poor substitute when the learner stops having to think.
The standard I want for plue is simple to state and demanding to meet: after using it, you should be better equipped to take the next step yourself. We are building toward that through context, attempts, feedback and review, and we intend to judge the work by what learners can demonstrate.
An education system should be able to answer the same question. Beyond what was submitted, beyond how long somebody sat there, what can they do now?
If you want to follow what we are building, join the plue waitlist. Joining the waitlist does not grant immediate product access.


