TGStat
TGStat
Type to search
Advanced channel search
  • flag English
    Site language
    flag Russian flag English flag Uzbek
  • Sign In
  • Catalog
    Channels and groups catalog Search for channels
    Add a channel/group
  • Ratings
    Rating of channels Rating of groups Posts rating
    Ratings of brands and people
  • Analytics
  • Search by posts
  • Telegram monitoring
Physical AI

13 Jul, 16:02

Open in Telegram Share Report

A useful way to think about robot training data.

Many teams building VLA models talk about data volume, but fewer explicitly decompose robot data into layers to discuss what each layer is actually useful for.

I spent three years scaling teleoperated data collection at Yango Robotics, from about 15,000 episodes a month to 400,000. Along the way, I learned that raw episodes / hours count can be often a misleading metric.
From my experience, data and labeling quality matter far more than people expect. A massive dataset only helps if it carries enough diversity, the right signal, and precise labels for the training stage you're targeting.
The real question devs should be asking is where your data sits on a crucial tradeoff: scale vs. contact fidelity.

Check out the video below for examples of the different data types, including the ones we make at Toloka.

Here is what that data spectrum looks like in practice:
Egocentric Video - great for scale, diversity, task context, and high-level planning. The limitation? It lacks ground-truth hand pose and a strong contact signal. You can estimate hand pose after the fact, but the quality rarely holds up for contact-rich manipulation on its own.

UMI-Style Data - costs more to collect, but gives you real 6DoF end-effector pose and a camera angle that actually sees contact. It sits right in the middle of the spectrum, some variants lean closer to egocentric collection, others closer to real teleop.

Real Robot Teleop - the most expensive and least scalable. It is also the only data that removes the sim-to-real and human-to-robot embodiment gaps entirely.

TLDR: If you want to teach a specific robot a specific task, the most direct path is usually to collect teleop data as close to that exact task, robot, environment, and action space as possible.
But if your goal is a generalist robot model, you probably need a layered approach: scalable egocentric data for breadth, higher-fidelity gripper data for contact, and real robot teleop for grounding.

Curious if you are thinking about this differently. Let's discuss in the comments.

https://www.linkedin.com/posts/v-toropov_a-useful-way-to-think-about-robot-training-ugcPost-7482415453257134080-rMB3/?utm_source=social_share_send&utm_medium=member_desktop_web&rcm=ACoAAAYuSGUB01rbMyTFNW4SkTf50dynCH7Luuw
A useful way to think about robot training data. Many teams building VLA models talk about data volume, but fewer explicitly decompose robot data into layers to discuss what each layer is actually… | Vasily Toropov
A useful way to think about robot training data. Many teams building VLA models talk about data volume, but fewer explicitly decompose robot data into layers to discuss what each layer is actually useful for. I spent three years scaling teleoperated data collection at Yango Robotics, from about 15...

4 0 0
Catalog
Channels and groups catalog Channels compilations Search for channels Add a channel/group
Ratings
Rating of Telegram channels Rating of Telegram groups Posts rating Ratings of brands and people
API
API statistics Search API of posts API Callback
Our channels
@TGStat @TGStat_Chat @telepulse @TGStatAPI
Read
Академия TGStat Telegram Research 2019 Telegram Research 2021 Telegram Research 2023
Contacts
Справочный центр Support Email Jobs
Miscellaneous
Terms and conditions Privacy policy Public offer
Our bots
@TGStat_Bot @SearcheeBot @TGAlertsBot @tg_analytics_bot @TGStatChatBot