Action Segmentation and Dense Captioning
The platform converts raw egocentric and robot video into structured data aligned with custom task structures. It identifies discrete, timestamped steps that can be mapped to user-provided verb lists and action hierarchies.
For language-conditioned robotic policies, the system generates natural-language descriptions that detail hand-object interactions, spatial relationships, and contextual scene dynamics without assembling separate transcription or framing tools.
- Detection of discrete, timestamped actions from egocentric video
- Custom taxonomy support for user-defined verbs and action hierarchies
- Rich natural-language descriptions covering hand-object interactions and spatial context
Quality Validation and Compliance Screening
To keep low-quality recordings out of model pipelines, the solution evaluates footage stability, framing, visual occlusion, and action clarity using configurable quality prompts.
The screening process also incorporates automated compliance checks that locate minors and personally identifiable information, including faces, license plates, badges, and documents, prior to delivery.
- Automated scoring for camera stability, framing, and occlusion
- Pre-review compliance screening for faces, documents, badges, and license plates
- Detection of minors to meet downstream privacy standards