Artificial intelligence (AI) is talked about everywhere these days, so it is natural to ask: can AI teach a child to speak? As someone who works with algorithms and data every day, my answer is — technology can do a lot, but not everything. And it is precisely knowing how these models work “from the inside” that helps you see their limits.
What technology genuinely can do
Modern tools really do lighten a great deal of the work:
- Quickly create, adapt and organise cards and photos.
- Structure the day — remind, measure, show what comes next.
- Reduce the routine workload of the educator and the parents, so more time is left for the child themselves.
That is real value. But notice: all of it is help for the adult, not the “teaching” of language to the child. The difference is fundamental.
Language is born in relationships
Language is not a file that can be “uploaded”. It is born in live relationships — with parents, an educator, the people closest to the child. A classic study by P. Kuhl and colleagues showed this beautifully. Nine-to-twelve-month-old American infants spent a few weeks interacting with a live person who spoke Mandarin, and they learned to distinguish Mandarin speech sounds. The infants who saw the very same content only on a screen, or heard only a recording, learned nothing (Kuhl, Tsao and Liu, 2003). Scientists call this the “social gate”: the brain opens up to language learning when a live person and shared attention are present (Kuhl, 2007).
A child learns words not from a “correct answer” but from a moment of joint attention — when they look at the same object together with an adult, share an emotion, wait for a response (Tomasello, 2003). The algorithm simply is not part of that relationship.
What research shows about screens
A 2020 review that pooled a large number of studies found: the more time a child spends in front of a screen, the weaker their language skills tend to be — yet quality content and watching together with an adult are associated with better outcomes (Madigan et al., 2020). The message is clear: what matters is not the screen itself but the person beside it. That is why the American Academy of Pediatrics emphasises not the “best app” but the adult's involvement (AAP, 2016).
Kalbeja takes this seriously: the app is designed to be a bridge between the child and a person, not one more screen the child is left alone with.
The “cold start” — why the algorithm is missing precisely your child
Here is the engineering explanation. AI models are trained on enormous but average data: the “average” child, the “average” environment, the “average” voice. Statistically, that is powerful. But your child is not an average. The link between the word “mum” and one particular face, between “outside” and the gate of your own yard — that is data no large model has, or ever can have. Specialists call this the “cold start” problem: the system knows nothing about a new, specific case.
That gap is filled not by a smarter algorithm but by the family — a real photo, a real voice, everyday context. The personalisation AI cannot “guess” is something a person creates in a few minutes.
Why familiar photos and a real voice beat “clever” algorithms
A child links a symbol to its meaning much faster when they see THEIR OWN environment and hear the voice of someone close to them. Research on AAC (augmentative and alternative communication) symbols shows: the more “transparent” a symbol is and the closer to the child's experience, the easier it is to understand (Light and Drager, 2007). A synthetic voice and generic pictures add extra cognitive load — the child first has to “decode” the symbol and only then use it. A real photo of mum, in mum's voice, needs no such step.
“Clever” does not always mean “right for the child”. That is exactly why Kalbeja deliberately builds on the human: the parents' voice, real photos and the educator's methodology. Kalbeja is not an AI product — and that is a conscious choice.
AAC does not hold back speech — what research shows
A common worry is: “if I give my child cards, maybe they will stop speaking?” The research is reassuring. A major review found that using AAC not only does not hold back speech but in many cases actually encourages it (Millar, Light and Schlosser, 2006). AAC is a bridge to language, not a substitute for it (Romski and Sevcik, 2005). Technology here serves the same goal — that the child starts to speak.
Where artificial intelligence can help — carefully
AI has a real place, but a clearly bounded one:
- Creating and organising content, translating, improving accessibility.
- Helping the adult work faster, so more time is left for the child.
On one condition — clear boundaries: privacy, human judgement and the child's interest first. Children's data is not “fuel” for models; Kalbeja runs on the device and feeds no AI with a child's information.
Artificial intelligence is a tool, not a teacher. Children are taught to speak by people; technology only helps them be together. As an engineer, I will put it simply: the best thing technology can do for language is step aside and leave room for the human.