# VideoDB leads a 100-hour action retrieval benchmark against AWS Nova and Twelve Labs

> A head-to-head retrieval benchmark, and what it means for robotics teams choosing what their next policy should learn from.

- Category: Benchmarks
- Published: 2026-10-06
- Authors: Sankalp Nagaonkar, Samuel Alexander
- Canonical: https://videodb.io/blog/a-faster-path-to-the-next-policy
- HTML: https://videodb.io/blog/a-faster-path-to-the-next-policy · Markdown: https://videodb.io/blog/a-faster-path-to-the-next-policy.md
- Tags: robotics, action segmentation, video retrieval, policy learning

---
On a 100-hour benchmark of 2,835 egocentric videos and 11,369 text queries, **VideoDB Action Segmentation + Search** returned a relevant moment in its top 10 results for **40.39%** of queries, about **5 points ahead of AWS Nova (35.40%)** and **7.5 points ahead of Twelve Labs Marengo 3.5 (32.87%)**. The gap widens when the result must match the action’s actual start and end: at temporal IoU ≥ 0.5, VideoDB reaches **20.49%**, about **2.05×** the next-best pipeline, and at 0.75 it reaches **12.83%**, about **3.81×**.

 <header class="figure-head"><h3 class="figure-title">Retrieval accuracy by pipeline on the 100-hour benchmark</h3><span class="figure-status"></span></header> <label for="result-k">Search depth <select id="result-k"><option value="1">Top 1</option><option value="5">Top 5</option><option value="10" selected>Top 10</option><option value="50">Top 50</option></select></label><label for="result-iou">Temporal match <select id="result-iou"><option value="0.5" selected>IoU ≥ 0.5</option><option value="0.75">IoU ≥ 0.75</option></select></label><a id="download-results" class="download" href="/assets/posts/a-faster-path-to-the-next-policy/results.csv" download="videodb-100h-retrieval-results.csv">Download data ↓</a> <section class="chart-panel" aria-labelledby="hit-question"><h4 id="hit-question">Find a relevant moment</h4><p class="chart-sub" id="hit-title">Hit@10 (%)</p>VideoDB<small>Action Segmentation + Search</small><span class="bar-fill" aria-hidden="true"></span><span class="bar-value">40.39</span>AWS Nova<span class="bar-fill" aria-hidden="true"></span><span class="bar-value">35.40</span>Twelve Labs<small>Marengo 3.5</small><span class="bar-fill" aria-hidden="true"></span><span class="bar-value">32.87</span>SigLIP2<span class="bar-fill" aria-hidden="true"></span><span class="bar-value">24.96</span>Qwen3-VL Embedding<span class="bar-fill" aria-hidden="true"></span><span class="bar-value">24.20</span>Cosmos-Embed<span class="bar-fill" aria-hidden="true"></span><span class="bar-value">23.78</span>InternVideo2<span class="bar-fill" aria-hidden="true"></span><span class="bar-value">19.50</span>PE-AV<span class="bar-fill" aria-hidden="true"></span><span class="bar-value">12.13</span><i></i><span style="left:0">0</span><span style="left:30.769%">20</span><span style="left:61.538%">40</span><span style="left:92.307%">60%</span></section><section class="chart-panel" aria-labelledby="temporal-question"><h4 id="temporal-question">Match the action’s interval</h4><p class="chart-sub" id="temporal-title">Hit@10 (%) · tIoU ≥ 0.5</p>VideoDB<small>Action Segmentation + Search</small><span class="bar-fill" aria-hidden="true"></span><span class="bar-value">20.49</span>AWS Nova<span class="bar-fill" aria-hidden="true"></span><span class="bar-value">10.01</span>Twelve Labs<small>Marengo 3.5</small><span class="bar-fill" aria-hidden="true"></span><span class="bar-value">9.61</span>SigLIP2<span class="bar-fill" aria-hidden="true"></span><span class="bar-value">6.22</span>Qwen3-VL Embedding<span class="bar-fill" aria-hidden="true"></span><span class="bar-value">5.30</span>Cosmos-Embed<span class="bar-fill" aria-hidden="true"></span><span class="bar-value">6.08</span>InternVideo2<span class="bar-fill" aria-hidden="true"></span><span class="bar-value">4.98</span>PE-AV<span class="bar-fill" aria-hidden="true"></span><span class="bar-value">2.44</span><i></i><span style="left:0">0</span><span style="left:30.769%">20</span><span style="left:61.538%">40</span><span style="left:92.307%">60%</span></section> <p class="sr-only" aria-live="polite" id="chart-announcement">Showing top 10 results with temporal IoU at least 0.5.</p>_<span class="fig-num">Figure 1.</span> VideoDB ranks first at top 1, 5 and 10, and on every temporal-match measure, in the evaluated configurations. Values show the percentage of queries with at least one matching result._

We also gave five open embedding models VideoDB’s action boundaries in place of fixed four-second windows. All five improved on temporally matched retrieval, by 1.71–3.57× at tIoU ≥ 0.5 (top 10), which suggests that the choice of temporal intervals is an important part of retrieval design alongside the embedding model.

Why this matters: before training the next robot policy, researchers have to find specific moments in hours of recordings, like a slip, a clean transfer or a recovery, and decide what the model should learn from. Below we explain how we measured retrieval, what we tested, and how this fits into that workflow.

 <header class="figure-head"><h3 class="figure-title">Example: from a failed run to a training selection</h3><span class="figure-status">Illustrative workflow</span></header>  <section class="hero-scene"><h4 class="phase"><span>01</span> Inspect the failure</h4><span class="episode-id">A</span><span>Slip during transfer</span><span class="time">05.17–15.63 s</span><span class="mini-timeline" aria-hidden="true"><i style="left:25.85%;width:52.300000000000004%"></i></span></section> <section><h4 class="phase"><span>02</span> Find related experience</h4><span class="episode-id">B</span><span>Clean transfer</span><span class="time">06.24–16.72 s</span><span class="mini-timeline" aria-hidden="true"><i style="left:31.2%;width:52.39999999999999%"></i></span><span class="episode-id">C</span><span>Recovery after a slip</span><span class="time">12.46–19.83 s</span><span class="mini-timeline" aria-hidden="true"><i style="left:62.3%;width:36.84999999999999%"></i></span><p class="phase-note">Compare the action and its outcome.</p></section> <section class="hero-handoff"><h4 class="phase"><span>03</span> Decide what to use</h4>Selection<span class="episode-id">B</span><span>Keep for comparison</span><span class="episode-id">C</span><span>Review the recovery</span><p>Source intervals and review decisions travel with the selection.</p></section> <p class="figure-takeaway">Which experience should inform the next iteration?</p>_<span class="fig-num">Figure 2.</span> A failed transfer (A) guides the search for a successful attempt (B) and a recovery (C), giving the researcher related experience to review for the next training selection._

Throughout this article, we use an illustrative cup-transfer task to follow that workflow: a robot loses its grip while moving a cup toward a tray, and a researcher locates the attempt and compares it with successful transfers and recoveries.

## Retrieve the action, not just the video

For this kind of investigation, a search result needs to capture enough of an action to explain what happened. A clip showing a cup in a gripper may be visually relevant, yet reveal little about whether the grasp was stable or why the transfer failed. The researcher needs a temporal unit that preserves the relationship between the attempt and its outcome, with a link back to the recording when more context is needed.

Action segmentation provides a way to organize video into those units by assigning descriptions and boundaries to individual actions. Our [revised WGO-Bench study](/blog/wgo-bench-action-annotations) evaluated how accurately VideoDB could identify action boundaries and label actions across 100 human and robot videos. VideoDB reached **28.93% semantic F1 at temporal IoU 0.5**, compared with **19.92%** for the public MacroData Refiner baseline. That study concerns the quality of the annotations; the next question is whether a retrieval pipeline can use action records to find relevant moments across a larger collection.

The boundaries of each result are important to that evaluation. In the cup example, a four-second window might capture the grasp without the slip, while another captures the drop without the approach. Either could return a relevant fragment, but the researcher would still have to reconstruct the sequence before deciding what it shows.

 <header class="figure-head"><h3 class="figure-title">How temporal IoU scores a returned clip</h3><span class="figure-status"></span></header> <p class="figure-context"><span class="episode-id">A</span> Slip during transfer </p> <span>Lift</span><span class="time">05.17 s</span><span>Transfer</span><span class="time">10.29 s</span><span>Slip</span><span class="time">15.63 s</span> <span>00.00 s</span><span>05.00 s</span><span>10.00 s</span><span>15.00 s</span><span>20.00 s</span> <span>Reference action</span><span>05.17–15.63 s</span>10.46 seconds <span>Returned interval</span><span id="candidate-times">07.42–11.86 s</span> <span><b id="intersection-duration">4.44 s</b> shared time</span><i aria-hidden="true">÷</i><span><b id="union-duration">10.46 s</b> union</span><i aria-hidden="true">=</i><span class="equation-result"><output id="iou-score" aria-live="polite">0.42</output> temporal IoU</span> <h4>Try a different interval</h4><button class="preset" data-start="7.42" data-end="11.86" aria-pressed="true">Too short</button><button class="preset" data-start="5.17" data-end="15.63" aria-pressed="false">Aligned</button><button class="preset" data-start="0" data-end="20" aria-pressed="false">Too wide</button> <p class="lens-note" id="lens-note">Only 4.44 of the action’s 10.46 seconds are shared. The interval misses part of the event.</p> <span id="check-05">0.5: not met</span><span id="check-075">0.75: not met</span> <details class="boundary-editor"><summary>Adjust the boundaries</summary><label class="control-label" for="interval-start">Start <output id="start-label" for="interval-start">07.42 s</output></label><input id="interval-start" type="range" min="0" max="19.99" step="0.01" value="7.42" aria-describedby="lens-note"><label class="control-label" for="interval-end">End <output id="end-label" for="interval-end">11.86 s</output></label><input id="interval-end" type="range" min="0.01" max="20" step="0.01" value="11.86" aria-describedby="lens-note"></details> _<span class="fig-num">Figure 3.</span> Shared time ÷ union measures temporal agreement. Too little context misses the event; too much adds unrelated time._

We measure this aspect of retrieval with **temporal intersection over union**, or tIoU: the shared time between a returned interval and a reference action, divided by the total time covered by either. A four-second fragment inside a ten-second action can reach at most 0.4 tIoU, even if every returned frame belongs to the action. Returning a much longer interval creates the opposite problem, preserving the event while adding unrelated activity that also lowers temporal agreement. The aim is to retrieve a well-bounded action while keeping the surrounding recording available for review.

## What 100 hours of retrieval tells us

To evaluate retrieval at a larger scale, we compared eight pipelines on **a 100-hour benchmark curated from [Lightwheel’s EgoStandard](https://huggingface.co/datasets/LightwheelAI/EgoStandard)**, comprising 2,835 videos, 11,369 text queries, and 17,967 reference intervals. The corpus captures human activity from an egocentric perspective, with reference intervals identifying individual actions within the recordings.

### How we compare the pipelines

Each pipeline searches the same videos with the same text queries and returns up to 50 timestamped candidates. What differs is how it divides the video, represents each interval, and ranks the results.

**VideoDB Action Segmentation + Search** generates action records with descriptions and temporal boundaries, then retrieves and ranks those records against the query. The approach connects semantic search to source-linked intervals, as described in our [search architecture report](https://labs.videodb.io/papers/search-over-the-visual-world.pdf).

**AWS Nova** uses Nova 2 Multimodal Embeddings through Amazon Bedrock Knowledge Bases, returning provider-managed, timestamped video chunks. **Twelve Labs** uses [Marengo 3.5](https://docs.twelvelabs.io/v1.3/docs/concepts/models/marengo/marengo-3-5) visual clip embeddings and the start and end times returned by its embedding service. Twelve Labs supports its own [temporal segmentation](https://docs.twelvelabs.io/v1.3/docs/guides/create-embeddings/at-scale/video). Marengo’s 512-dimensional embeddings are normalized and searched locally with FAISS.

For the local comparison, we selected five prominent open embedding models: [SigLIP2](https://huggingface.co/google/siglip2-so400m-patch14-384), [Qwen3-VL Embedding](https://huggingface.co/Qwen/Qwen3-VL-Embedding-2B), [Cosmos-Embed](https://huggingface.co/nvidia/Cosmos-Embed1-336p), [InternVideo2](https://github.com/OpenGVLab/InternVideo/tree/main/InternVideo2/multi_modality), and [PE-AV](https://huggingface.co/facebook/pe-av-large). We encode **four-second windows sampled at 2 FPS**, giving eight frame samples per full window, with each model’s native preprocessing. We use **FAISS as the local vector index** to store normalized embeddings.

### What counts as a successful result

We score the first K results in two ways, both requiring a match in the **same source video**:

- **Hit@K** uses a minimum-overlap threshold. At least one returned interval must overlap a reference action by **min(0.5 seconds, 50% of the reference duration)**. For an action lasting ten seconds, just half a second of overlap is enough; for a 0.4-second action, the threshold is 0.2 seconds. This is a permissive overlap check, so it can count a fragment without capturing the full action.
- **Hit@K at temporal IoU 0.5 or 0.75** uses a stricter temporal-overlap threshold. At least one result must meet the specified ratio of shared time to total covered time. Both where the interval starts and where it ends affect this score: missing part of the action or including extra footage outside it lowers temporal IoU. A higher threshold demands closer temporal agreement.

Both metrics report the percentage of queries with at least one qualifying result. The difference is how much temporal agreement is needed for a result to count as a match.

VideoDB leads Hit@1, Hit@5, and Hit@10, with **40.39% at Hit@10**, compared with **35.40%** for AWS Nova and **32.87%** for Twelve Labs. Its lead is larger on temporally matched retrieval: at Hit@10 (tIoU ≥ 0.5), VideoDB reaches **20.49%**, about **2.05×** the next-best pipeline’s 10.01%. At the stricter 0.75 threshold, it reaches **12.83%**, about **3.81×** the next-best pipeline’s 3.37%, and leads every reported IoU-qualified column.

AWS Nova and Twelve Labs achieve higher Hit@50 under the permissive overlap rule. VideoDB leads at every reported retrieval depth on Hit@K at both tIoU thresholds, which require the retrieved interval to match the reference action more closely.

For researchers reviewing search results, early ranking and temporal agreement are useful properties because they determine what appears first and how much of the action it contains. The comparison evaluates each complete pipeline, including its representation, temporal units, and ranking.

## A search result is not a training decision

The distinction between relevance and suitability becomes clear when we return to the cup transfer. A successful attempt could help a researcher compare approaches and grasps, a failed attempt could support diagnosis, and a recovery could be useful for a different learning objective. All three may belong in the search results, but deciding which should enter training requires inspecting the sequence and understanding the supervision available in its source.

 <header class="figure-head"><h3 class="figure-title">Reviewing search results before they go into training</h3><span class="figure-status"></span></header> <p class="selection-question">“Find comparable transfers, including slips and recoveries.”</p> <button type="button" class="candidate" data-episode="A" data-start="5.17" data-end="15.63" aria-pressed="false" aria-controls="selection-record"><span class="candidate-info"><span class="episode-label"><span class="episode-id">A</span><span>Slip during transfer</span></span><span class="time">05.17–15.63 s</span><span class="mini-timeline" aria-hidden="true"><i style="left:25.85%;width:52.300000000000004%"></i></span><span class="candidate-use" data-status="hold">Hold for diagnosis</span></span><span class="candidate-arrow" aria-hidden="true">↗</span></button><button type="button" class="candidate" data-episode="B" data-start="6.24" data-end="16.72" aria-pressed="true" aria-controls="selection-record"><span class="candidate-info"><span class="episode-label"><span class="episode-id">B</span><span>Clean transfer</span></span><span class="time">06.24–16.72 s</span><span class="mini-timeline" aria-hidden="true"><i style="left:31.2%;width:52.39999999999999%"></i></span><span class="candidate-use" data-status="keep">Keep for comparison</span></span><span class="candidate-arrow" aria-hidden="true">↗</span></button><button type="button" class="candidate" data-episode="C" data-start="12.46" data-end="19.83" aria-pressed="false" aria-controls="selection-record"><span class="candidate-info"><span class="episode-label"><span class="episode-id">C</span><span>Recovery after a slip</span></span><span class="time">12.46–19.83 s</span><span class="mini-timeline" aria-hidden="true"><i style="left:62.3%;width:36.84999999999999%"></i></span><span class="candidate-use" data-status="review">Review the recovery</span></span><span class="candidate-arrow" aria-hidden="true">↗</span></button><p class="interaction-hint">Select an episode to inspect its proposed record.</p> <section class="selection-manifest" id="selection-record" aria-labelledby="selection-record-title" aria-live="polite" data-episode="B" data-start="6.24" data-end="16.72"><header><h4 id="selection-record-title">Selection record</h4></header> <span class="episode-id" id="record-id">B</span><strong id="record-name">Clean transfer</strong> <dl><dt>Source</dt><dd id="record-source">Episode B · original recording</dd><dt>Interval</dt><dd id="record-interval">06.24–16.72 s</dd><dt>Decision</dt><dd id="record-decision">Keep for comparison</dd><dt>Purpose</dt><dd id="record-purpose">Compare approach, grasp, and outcome</dd></dl> <p class="record-next" id="record-next">Confirm compatible supervision before adding it to training.</p><p class="record-provenance">Retain source identity, timestamps, and query, label, and review versions.</p> </section>_<span class="fig-num">Figure 4.</span> Search proposes candidates. Review decides their role. The same episode IDs and source intervals persist into the proposed selection record._

Each selected interval should carry enough context to explain where it came from, why it was chosen, and how it is intended to be used. That includes the source episode, timestamps, the origin of its labels, and the reviewer’s decision. Where compatible robot observations, actions, and state are available, the selection should preserve their relationship to the video and the control schema needed to interpret them. This makes it possible to revisit a selection when an experiment succeeds, fails, or reveals an unexpected regression.

The source of the experience also constrains how it can be used. Human video can support visual representations, task structure, or semantic supervision, but it does not automatically provide robot control targets. Similarly, a failed action may be informative for diagnosis without being an appropriate behavior-cloning target. Retrieval should help researchers make these distinctions rather than treat every relevant clip as interchangeable training data.

Existing research illustrates several ways action structure can contribute to learning. [π0.5](https://arxiv.org/abs/2504.16054) combines semantic subtask supervision with low-level action learning, while [SARM](https://arxiv.org/abs/2509.25358v4) uses subtask annotations for stage and progress supervision, then learned rewards to filter and reweight demonstrations. [COLLAGE](https://arxiv.org/abs/2508.01131) approaches demonstration selection through multiple retrieval cues and task-specific weighting. These methods use different learning objectives, but each makes the role of the selected or annotated experience explicit.

Search could also help investigate changes between policy versions. If a new checkpoint appears to drop cups more often, researchers could retrieve attempts linked to each version and compare grasps, slips, and recoveries under similar conditions. This helps distinguish a possible policy regression from changes in objects or hardware, while consistent outcome labels and attempt counts provide the basis for measuring it.

That comparison could then guide the next data decision: select successful grasps under the affected conditions, review recovery examples, or collect new demonstrations if those cases are missing. More generally, when the available recordings do not cover the behavior a researcher needs to study, the search can help turn a broad request for more data into a specific collection target.

## Do VideoDB’s action boundaries help other embedding models?

On the same 100-hour benchmark, we also compared five embedding models using the action boundaries from VideoDB, which leads the temporal-IoU retrieval metrics above. Rather than dividing the videos into four-second windows, we gave Cosmos-Embed, SigLIP2, PE-AV, Qwen3-VL Embedding, and InternVideo2 the same action intervals and let each model encode the original frames within them. Only the segmentation is shared; each model computes its own embeddings. Comparing these with each model’s four-second-window scores shows how much the choice of interval affects retrieval.

 <header class="figure-head"><h3 class="figure-title">Test setup: five embedding models on VideoDB action intervals</h3><span class="figure-status"></span></header>  <h4>VideoDB action boundaries</h4><span style="left:0.0%;width:25.85%"><i></i></span><span style="left:25.85%;width:52.300000000000004%"><i></i></span><span style="left:78.15%;width:21.849999999999998%"><i></i></span><p class="shared-input-note">Same intervals · Original frames</p> <span class="paired-arrow" aria-hidden="true"></span> <ul class="shared-encoders"><li>Cosmos-Embed</li><li>SigLIP2</li><li>PE-AV</li><li>Qwen3-VL Embedding</li><li>InternVideo2</li></ul> <span class="paired-arrow" aria-hidden="true"></span> <strong>Compare retrieval</strong><span>Hit@K · Temporal IoU</span> _<span class="fig-num">Figure 5.</span> All five embedding models encode the same action intervals from the 100-hour benchmark._

The comparison uses the same corpus, queries, and reference intervals. Results for each model are below.

<section class="wide-figure draft-results" aria-labelledby="paired-results-title"><header class="figure-head"><h3 class="figure-title" id="paired-results-title">Four-second windows vs. VideoDB action intervals, by model</h3></header><p id="paired-results-note">Values are query success percentages.</p><label for="paired-k">Search depth <select id="paired-k"><option value="1">Top 1</option><option value="5">Top 5</option><option value="10" selected>Top 10</option><option value="50">Top 50</option></select></label><table class="paired-results-table" aria-describedby="paired-results-note"><caption class="sr-only">Each cell shows four-second windows followed by VideoDB action intervals.</caption><thead><tr><th scope="col">Encoder</th><th scope="col"><span class="paired-metric">Hit@10</span></th><th scope="col"><span class="paired-metric">Hit@10</span><br>tIoU ≥ 0.5</th><th scope="col"><span class="paired-metric">Hit@10</span><br>tIoU ≥ 0.75</th></tr></thead><tbody><tr><th scope="row">Cosmos-Embed</th><td><span class="paired-before">23.78</span> → <strong class="paired-after">24.94</strong></td><td><span class="paired-before">6.08</span> → <strong class="paired-after">12.06</strong></td><td><span class="paired-before">2.08</span> → <strong class="paired-after">7.50</strong></td></tr><tr><th scope="row">SigLIP2</th><td><span class="paired-before">24.96</span> → <strong class="paired-after">26.91</strong></td><td><span class="paired-before">6.22</span> → <strong class="paired-after">12.25</strong></td><td><span class="paired-before">1.90</span> → <strong class="paired-after">6.93</strong></td></tr><tr><th scope="row">PE-AV</th><td><span class="paired-before">12.13</span> → <strong class="paired-after">16.47</strong></td><td><span class="paired-before">2.44</span> → <strong class="paired-after">8.70</strong></td><td><span class="paired-before">0.82</span> → <strong class="paired-after">5.19</strong></td></tr><tr><th scope="row">Qwen3-VL Embedding</th><td><span class="paired-before">24.20</span> → <strong class="paired-after">26.58</strong></td><td><span class="paired-before">5.30</span> → <strong class="paired-after">12.75</strong></td><td><span class="paired-before">1.79</span> → <strong class="paired-after">7.88</strong></td></tr><tr><th scope="row">InternVideo2</th><td><span class="paired-before">19.50</span> → <strong class="paired-after">20.02</strong></td><td><span class="paired-before">4.98</span> → <strong class="paired-after">8.51</strong></td><td><span class="paired-before">1.61</span> → <strong class="paired-after">5.07</strong></td></tr></tbody></table><p class="sr-only" id="paired-announcement" aria-live="polite">Showing top 10 results for five embedding models.</p></section>

Across all five tested embedding models, replacing four-second windows with VideoDB action intervals improves temporally matched retrieval at every reported depth. At top 10, success rates increase by **1.71–3.57× at tIoU ≥ 0.5** and **3.15–6.33× at tIoU ≥ 0.75**. The improvement extends across encoders, supporting the hypothesis that the choice of temporal intervals is an important part of retrieval design alongside the embedding model.

The improvement is most consistent when retrieval is evaluated against the action’s temporal boundaries. Hit@K counts a result once it meets a minimum-overlap threshold, even if it captures only a small part of the action. Temporal IoU demands closer agreement with the reference interval. Under this stricter measure, VideoDB action intervals improve retrieval across all five encoders at every reported depth, showing a consistent benefit in retrieving intervals that better match the actions being searched for.

## Measure progress at the next model

 <header class="figure-head"><h3 class="figure-title">Where search and review fit in the training loop</h3><span class="figure-status"></span></header>  <section class="cycle-stage"><h4>Observe the current policy</h4><span class="episode-id">A</span><span>Find the attempt.<br>Inspect what happened.</span></section> <span>Search<br> + review</span><i>→</i> <section class="cycle-stage cycle-curation"><h4>Build a training selection</h4><header><span>Source-linked records</span></header><span class="episode-id">B</span><span>Clean transfer</span><span class="time">06.24–16.72 s</span><span class="mini-timeline" aria-hidden="true"><i style="left:31.2%;width:52.39999999999999%"></i></span><span class="episode-id">C</span><span>Recovery after a slip</span><span class="time">12.46–19.83 s</span><span class="mini-timeline" aria-hidden="true"><i style="left:62.3%;width:36.84999999999999%"></i></span><p>Review suitability for training</p></section> <span>Curate<br> + retrain</span><i>→</i> <section class="cycle-stage"><h4>Evaluate the next policy</h4><p>Test the target behavior<br>and check for regressions</p></section> <p><span>New runs reveal what to investigate next</span></p>_<span class="fig-num">Figure 6.</span> Find useful experience, review it for training, and evaluate the next policy on the behavior you want to improve._

Our broader goal is to make the work between model versions faster by reducing the effort it takes to turn recorded robot experience into useful training selections. VideoDB Action Segmentation + Search helps researchers locate relevant actions across recordings, inspect their context, and decide which examples belong in the next experiment. We want researchers to spend less time finding evidence and more time testing what the next policy should learn, so the experience they have already collected becomes a better foundation for the next model.

