🔄 Russian Mathematicians Cut AI Computing Costs 20x
Startup Mostik AI has developed a way for AI models to communicate directly by sharing their internal states rather than passing text back and forth. The approach lets a large model do the computationally intensive work of processing a task, while a smaller, cheaper model generates the response.
⚙️ Mostik AI CEO Sasha Malysheva told @hiaimediaen how it works:
Multi-model systems are already used in AI agents to distribute work between models with different capabilities and costs. One model might plan a series of actions, for example, while another executes them. But they still typically communicate through text—one model generates an output, and the next one reads it.
When processing each token, a model builds numerical vectors that make up its hidden state. This internal state contains the context, information about how the model understands the task, and what a possible answer might look like.
According to Mostik, producing a single token involves more than 100 hidden vectors, or about 1 million numbers—roughly 2 MB of internal state. Yet only one selected token leaves the model, equivalent to about 17 bits. "Systems built to think in thousands of dimensions are effectively talking to each other through a keyhole," the developers write.
Mostik ("little bridge" in Russian) has built, fittingly, a small "bridge" that translates the hidden states of one model into a representation another model can work with directly. Malysheva explains it with the example of a student asking for help with differential equations before an exam:
For the demonstration, the researchers connected GLM-5.2, a 753 billion-parameter model that can require up to 1.5 TB of memory to run, with Qwen-3.5, a 4 billion-parameter model compact enough to run on a smartphone. According to Malysheva, the hybrid system delivered an average of 80% of the larger model's quality at roughly one-twentieth of the compute cost.
💡 Why Does This Matter?
With the bridge, developers could use large models only for the parts of a task that actually require them, while handing the rest of the work to cheaper, compact models. For users, that could mean better answers at a lower cost—without having to constantly think about which model they are talking to.
🎓 A Dream Team
The Mostik team was assembled in four months. The startup is now based in California and has 15 employees, 12 of whom hold PhDs. Its first employees were Malysheva's former classmates at HSE University in St. Petersburg—they have been building projects together for more than 12 years. The rest were recruited through friends, with each person brought in to cover a specific area of expertise.
The project's chief scientist is Stanislav Smirnov, a professor at the University of Geneva and winner of the 2010 Fields Medal. Finding common ground between two AI models, he says, is surprisingly difficult. "There seems to be no appropriate mathematical language yet."
Smirnov believes that deeper mathematical analysis could eventually reveal commonalities in how AI models and humans reason about difficult problems.
@hiaimediaen
Startup Mostik AI has developed a way for AI models to communicate directly by sharing their internal states rather than passing text back and forth. The approach lets a large model do the computationally intensive work of processing a task, while a smaller, cheaper model generates the response.
⚙️ Mostik AI CEO Sasha Malysheva told @hiaimediaen how it works:
Multi-model systems are already used in AI agents to distribute work between models with different capabilities and costs. One model might plan a series of actions, for example, while another executes them. But they still typically communicate through text—one model generates an output, and the next one reads it.
When processing each token, a model builds numerical vectors that make up its hidden state. This internal state contains the context, information about how the model understands the task, and what a possible answer might look like.
According to Mostik, producing a single token involves more than 100 hidden vectors, or about 1 million numbers—roughly 2 MB of internal state. Yet only one selected token leaves the model, equivalent to about 17 bits. "Systems built to think in thousands of dimensions are effectively talking to each other through a keyhole," the developers write.
Mostik ("little bridge" in Russian) has built, fittingly, a small "bridge" that translates the hidden states of one model into a representation another model can work with directly. Malysheva explains it with the example of a student asking for help with differential equations before an exam:
"A large model comes up with a good analogy that will help the student remember the concept and passes these 'thoughts' to a smaller model, which then writes out the solution and develops the analogy. The second model benefits from the work the first one has already done, instead of having to come up with something on its own and getting a worse result by definition—because it is a less capable model."
For the demonstration, the researchers connected GLM-5.2, a 753 billion-parameter model that can require up to 1.5 TB of memory to run, with Qwen-3.5, a 4 billion-parameter model compact enough to run on a smartphone. According to Malysheva, the hybrid system delivered an average of 80% of the larger model's quality at roughly one-twentieth of the compute cost.
💡 Why Does This Matter?
With the bridge, developers could use large models only for the parts of a task that actually require them, while handing the rest of the work to cheaper, compact models. For users, that could mean better answers at a lower cost—without having to constantly think about which model they are talking to.
🎓 A Dream Team
The Mostik team was assembled in four months. The startup is now based in California and has 15 employees, 12 of whom hold PhDs. Its first employees were Malysheva's former classmates at HSE University in St. Petersburg—they have been building projects together for more than 12 years. The rest were recruited through friends, with each person brought in to cover a specific area of expertise.
The project's chief scientist is Stanislav Smirnov, a professor at the University of Geneva and winner of the 2010 Fields Medal. Finding common ground between two AI models, he says, is surprisingly difficult. "There seems to be no appropriate mathematical language yet."
Smirnov believes that deeper mathematical analysis could eventually reveal commonalities in how AI models and humans reason about difficult problems.
@hiaimediaen