All posts
Benchmarks Sankalp Nagaonkar, Samuel Alexander

VideoDB leads a 100-hour action retrieval benchmark against AWS Nova and Twelve Labs

A head-to-head retrieval benchmark, and what it means for robotics teams choosing what their next policy should learn from.

On a 100-hour benchmark of 2,835 egocentric videos and 11,369 text queries, VideoDB Action Segmentation + Search returned a relevant moment in its top 10 results for 40.39% of queries, about 5 points ahead of AWS Nova (35.40%) and 7.5 points ahead of Twelve Labs Marengo 3.5 (32.87%). The gap widens when the result must match the action’s actual start and end: at temporal IoU ≥ 0.5, VideoDB reaches 20.49%, about 2.05× the next-best pipeline, and at 0.75 it reaches 12.83%, about 3.81×.

Retrieval accuracy by pipeline on the 100-hour benchmark

Download data ↓

Find a relevant moment

Hit@10 (%)

VideoDBAction Segmentation + Search
40.39
AWS Nova
35.40
Twelve LabsMarengo 3.5
32.87
SigLIP2
24.96
Qwen3-VL Embedding
24.20
Cosmos-Embed
23.78
InternVideo2
19.50
PE-AV
12.13

Match the action’s interval

Hit@10 (%) · tIoU ≥ 0.5

VideoDBAction Segmentation + Search
20.49
AWS Nova
10.01
Twelve LabsMarengo 3.5
9.61
SigLIP2
6.22
Qwen3-VL Embedding
5.30
Cosmos-Embed
6.08
InternVideo2
4.98
PE-AV
2.44

Showing top 10 results with temporal IoU at least 0.5.

Figure 1. VideoDB ranks first at top 1, 5 and 10, and on every temporal-match measure, in the evaluated configurations. Values show the percentage of queries with at least one matching result.

We also gave five open embedding models VideoDB’s action boundaries in place of fixed four-second windows. All five improved on temporally matched retrieval, by 1.71–3.57× at tIoU ≥ 0.5 (top 10), which suggests that the choice of temporal intervals is an important part of retrieval design alongside the embedding model.

Why this matters: before training the next robot policy, researchers have to find specific moments in hours of recordings, like a slip, a clean transfer or a recovery, and decide what the model should learn from. Below we explain how we measured retrieval, what we tested, and how this fits into that workflow.

Example: from a failed run to a training selection

Illustrative workflow

01 Inspect the failure

Schematic robot arm losing an orange cup before it reaches a tray.
ASlip during transfer
05.17–15.63 s

02 Find related experience

Compare the action and its outcome.

03 Decide what to use

Selection
BKeep for comparison
CReview the recovery

Source intervals and review decisions travel with the selection.

Which experience should inform the next iteration?

Figure 2. A failed transfer (A) guides the search for a successful attempt (B) and a recovery (C), giving the researcher related experience to review for the next training selection.

Throughout this article, we use an illustrative cup-transfer task to follow that workflow: a robot loses its grip while moving a cup toward a tray, and a researcher locates the attempt and compares it with successful transfers and recoveries.

Retrieve the action, not just the video

For this kind of investigation, a search result needs to capture enough of an action to explain what happened. A clip showing a cup in a gripper may be visually relevant, yet reveal little about whether the grasp was stable or why the transfer failed. The researcher needs a temporal unit that preserves the relationship between the attempt and its outcome, with a link back to the recording when more context is needed.

Action segmentation provides a way to organize video into those units by assigning descriptions and boundaries to individual actions. Our revised WGO-Bench study evaluated how accurately VideoDB could identify action boundaries and label actions across 100 human and robot videos. VideoDB reached 28.93% semantic F1 at temporal IoU 0.5, compared with 19.92% for the public MacroData Refiner baseline. That study concerns the quality of the annotations; the next question is whether a retrieval pipeline can use action records to find relevant moments across a larger collection.

The boundaries of each result are important to that evaluation. In the cup example, a four-second window might capture the grasp without the slip, while another captures the drop without the approach. Either could return a relevant fragment, but the researcher would still have to reconstruct the sequence before deciding what it shows.

How temporal IoU scores a returned clip

A Slip during transfer

Lift05.17 s
Transfer10.29 s
Slip15.63 s
Reference action05.17–15.63 s
10.46 seconds
Returned interval07.42–11.86 s
4.44 s shared time10.46 s union0.42 temporal IoU

Try a different interval

Only 4.44 of the action’s 10.46 seconds are shared. The interval misses part of the event.

0.5: not met0.75: not met
Adjust the boundaries
Figure 3. Shared time ÷ union measures temporal agreement. Too little context misses the event; too much adds unrelated time.

We measure this aspect of retrieval with temporal intersection over union, or tIoU: the shared time between a returned interval and a reference action, divided by the total time covered by either. A four-second fragment inside a ten-second action can reach at most 0.4 tIoU, even if every returned frame belongs to the action. Returning a much longer interval creates the opposite problem, preserving the event while adding unrelated activity that also lowers temporal agreement. The aim is to retrieve a well-bounded action while keeping the surrounding recording available for review.

What 100 hours of retrieval tells us

To evaluate retrieval at a larger scale, we compared eight pipelines on a 100-hour benchmark curated from Lightwheel’s EgoStandard, comprising 2,835 videos, 11,369 text queries, and 17,967 reference intervals. The corpus captures human activity from an egocentric perspective, with reference intervals identifying individual actions within the recordings.

How we compare the pipelines

Each pipeline searches the same videos with the same text queries and returns up to 50 timestamped candidates. What differs is how it divides the video, represents each interval, and ranks the results.

VideoDB Action Segmentation + Search generates action records with descriptions and temporal boundaries, then retrieves and ranks those records against the query. The approach connects semantic search to source-linked intervals, as described in our search architecture report.

AWS Nova uses Nova 2 Multimodal Embeddings through Amazon Bedrock Knowledge Bases, returning provider-managed, timestamped video chunks. Twelve Labs uses Marengo 3.5 visual clip embeddings and the start and end times returned by its embedding service. Twelve Labs supports its own temporal segmentation. Marengo’s 512-dimensional embeddings are normalized and searched locally with FAISS.

For the local comparison, we selected five prominent open embedding models: SigLIP2, Qwen3-VL Embedding, Cosmos-Embed, InternVideo2, and PE-AV. We encode four-second windows sampled at 2 FPS, giving eight frame samples per full window, with each model’s native preprocessing. We use FAISS as the local vector index to store normalized embeddings.

What counts as a successful result

We score the first K results in two ways, both requiring a match in the same source video:

  • Hit@K uses a minimum-overlap threshold. At least one returned interval must overlap a reference action by min(0.5 seconds, 50% of the reference duration). For an action lasting ten seconds, just half a second of overlap is enough; for a 0.4-second action, the threshold is 0.2 seconds. This is a permissive overlap check, so it can count a fragment without capturing the full action.
  • Hit@K at temporal IoU 0.5 or 0.75 uses a stricter temporal-overlap threshold. At least one result must meet the specified ratio of shared time to total covered time. Both where the interval starts and where it ends affect this score: missing part of the action or including extra footage outside it lowers temporal IoU. A higher threshold demands closer temporal agreement.

Both metrics report the percentage of queries with at least one qualifying result. The difference is how much temporal agreement is needed for a result to count as a match.

VideoDB leads Hit@1, Hit@5, and Hit@10, with 40.39% at Hit@10, compared with 35.40% for AWS Nova and 32.87% for Twelve Labs. Its lead is larger on temporally matched retrieval: at Hit@10 (tIoU ≥ 0.5), VideoDB reaches 20.49%, about 2.05× the next-best pipeline’s 10.01%. At the stricter 0.75 threshold, it reaches 12.83%, about 3.81× the next-best pipeline’s 3.37%, and leads every reported IoU-qualified column.

AWS Nova and Twelve Labs achieve higher Hit@50 under the permissive overlap rule. VideoDB leads at every reported retrieval depth on Hit@K at both tIoU thresholds, which require the retrieved interval to match the reference action more closely.

For researchers reviewing search results, early ranking and temporal agreement are useful properties because they determine what appears first and how much of the action it contains. The comparison evaluates each complete pipeline, including its representation, temporal units, and ranking.

A search result is not a training decision

The distinction between relevance and suitability becomes clear when we return to the cup transfer. A successful attempt could help a researcher compare approaches and grasps, a failed attempt could support diagnosis, and a recovery could be useful for a different learning objective. All three may belong in the search results, but deciding which should enter training requires inspecting the sequence and understanding the supervision available in its source.

Reviewing search results before they go into training

“Find comparable transfers, including slips and recoveries.”

Select an episode to inspect its proposed record.

Selection record

BClean transfer
Source
Episode B · original recording
Interval
06.24–16.72 s
Decision
Keep for comparison
Purpose
Compare approach, grasp, and outcome

Confirm compatible supervision before adding it to training.

Retain source identity, timestamps, and query, label, and review versions.

Figure 4. Search proposes candidates. Review decides their role. The same episode IDs and source intervals persist into the proposed selection record.

Each selected interval should carry enough context to explain where it came from, why it was chosen, and how it is intended to be used. That includes the source episode, timestamps, the origin of its labels, and the reviewer’s decision. Where compatible robot observations, actions, and state are available, the selection should preserve their relationship to the video and the control schema needed to interpret them. This makes it possible to revisit a selection when an experiment succeeds, fails, or reveals an unexpected regression.

The source of the experience also constrains how it can be used. Human video can support visual representations, task structure, or semantic supervision, but it does not automatically provide robot control targets. Similarly, a failed action may be informative for diagnosis without being an appropriate behavior-cloning target. Retrieval should help researchers make these distinctions rather than treat every relevant clip as interchangeable training data.

Existing research illustrates several ways action structure can contribute to learning. π0.5 combines semantic subtask supervision with low-level action learning, while SARM uses subtask annotations for stage and progress supervision, then learned rewards to filter and reweight demonstrations. COLLAGE approaches demonstration selection through multiple retrieval cues and task-specific weighting. These methods use different learning objectives, but each makes the role of the selected or annotated experience explicit.

Search could also help investigate changes between policy versions. If a new checkpoint appears to drop cups more often, researchers could retrieve attempts linked to each version and compare grasps, slips, and recoveries under similar conditions. This helps distinguish a possible policy regression from changes in objects or hardware, while consistent outcome labels and attempt counts provide the basis for measuring it.

That comparison could then guide the next data decision: select successful grasps under the affected conditions, review recovery examples, or collect new demonstrations if those cases are missing. More generally, when the available recordings do not cover the behavior a researcher needs to study, the search can help turn a broad request for more data into a specific collection target.

Do VideoDB’s action boundaries help other embedding models?

On the same 100-hour benchmark, we also compared five embedding models using the action boundaries from VideoDB, which leads the temporal-IoU retrieval metrics above. Rather than dividing the videos into four-second windows, we gave Cosmos-Embed, SigLIP2, PE-AV, Qwen3-VL Embedding, and InternVideo2 the same action intervals and let each model encode the original frames within them. Only the segmentation is shared; each model computes its own embeddings. Comparing these with each model’s four-second-window scores shows how much the choice of interval affects retrieval.

Test setup: five embedding models on VideoDB action intervals

VideoDB action boundaries

Same intervals · Original frames

  • Cosmos-Embed
  • SigLIP2
  • PE-AV
  • Qwen3-VL Embedding
  • InternVideo2
Compare retrievalHit@K · Temporal IoU
Figure 5. All five embedding models encode the same action intervals from the 100-hour benchmark.

The comparison uses the same corpus, queries, and reference intervals. Results for each model are below.

Four-second windows vs. VideoDB action intervals, by model

Values are query success percentages.

Each cell shows four-second windows followed by VideoDB action intervals.
EncoderHit@10Hit@10
tIoU ≥ 0.5
Hit@10
tIoU ≥ 0.75
Cosmos-Embed23.78 → 24.946.08 → 12.062.08 → 7.50
SigLIP224.96 → 26.916.22 → 12.251.90 → 6.93
PE-AV12.13 → 16.472.44 → 8.700.82 → 5.19
Qwen3-VL Embedding24.20 → 26.585.30 → 12.751.79 → 7.88
InternVideo219.50 → 20.024.98 → 8.511.61 → 5.07

Showing top 10 results for five embedding models.

Across all five tested embedding models, replacing four-second windows with VideoDB action intervals improves temporally matched retrieval at every reported depth. At top 10, success rates increase by 1.71–3.57× at tIoU ≥ 0.5 and 3.15–6.33× at tIoU ≥ 0.75. The improvement extends across encoders, supporting the hypothesis that the choice of temporal intervals is an important part of retrieval design alongside the embedding model.

The improvement is most consistent when retrieval is evaluated against the action’s temporal boundaries. Hit@K counts a result once it meets a minimum-overlap threshold, even if it captures only a small part of the action. Temporal IoU demands closer agreement with the reference interval. Under this stricter measure, VideoDB action intervals improve retrieval across all five encoders at every reported depth, showing a consistent benefit in retrieving intervals that better match the actions being searched for.

Measure progress at the next model

Where search and review fit in the training loop

Observe the current policy

The current policy’s recorded cup-transfer attempt.
AFind the attempt.
Inspect what happened.

Build a training selection

Source-linked records
BClean transfer
06.24–16.72 s
CRecovery after a slip
12.46–19.83 s

Review suitability for training

Evaluate the next policy

A new policy attempt under evaluation; its outcome is not shown.

Test the target behavior
and check for regressions

New runs reveal what to investigate next

Figure 6. Find useful experience, review it for training, and evaluate the next policy on the behavior you want to improve.

Our broader goal is to make the work between model versions faster by reducing the effort it takes to turn recorded robot experience into useful training selections. VideoDB Action Segmentation + Search helps researchers locate relevant actions across recordings, inspect their context, and decide which examples belong in the next experiment. We want researchers to spend less time finding evidence and more time testing what the next policy should learn, so the experience they have already collected becomes a better foundation for the next model.

Machine

https://videodb.io/blog/a-faster-path-to-the-next-policy.mdOpen the file