Evaluating VLAs and Robot Foundation Models
Hosted by Weights & Biases by CoreWeave and Cyber Valley
As physical AI teams move from simulation to deployment, the question is no longer just whether a system works but whether it performs reliably and safely in the real world. Join Weights & Biases by CoreWeave for a deep dive into how robotics teams evaluate checkpoints before deploying them in real-world systems.
Moderated by Hans Ramsl, Senior Accounts Solutions Architect at Weights & Biases by CoreWeave, the session will feature a guest speaker from Polybot and explore:
- Sim-to-real gaps: What does a strong simulation benchmark actually predict about a deployed policy?
- Beyond success rates: Why do teams track trajectory traces, intervention rates, and near-miss taxonomies?
- The release decision: How do teams determine whether a checkpoint is safe to deploy in a customer’s cell?
Key takeaways from the session include:
● Shared vocabulary for the signals that matter in addition to success rates
● Practical ways to turn evaluation results into actions on new checkpoints
● Opportunity to network with peers across the Cyber Valley network and the Weights and Biases by CoreWeave team
When
September 17, 17:00–19:30
Where
Don’t miss out, register via Eventbrite today!