<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>Research &amp; Blog · VideoDB</title>
    <link>https://videodb.io/blog</link>
    <atom:link href="https://videodb.io/rss.xml" rel="self" type="application/rss+xml" />
    <description>Papers, benchmark results, engineering notes, essays and apps from the VideoDB lab.</description>
    <language>en-us</language>
    <lastBuildDate>Tue, 06 Oct 2026 00:00:00 GMT</lastBuildDate>
    <item>
      <title>VideoDB leads a 100-hour action retrieval benchmark against AWS Nova and Twelve Labs</title>
      <link>https://videodb.io/blog/a-faster-path-to-the-next-policy</link>
      <guid isPermaLink="true">https://videodb.io/blog/a-faster-path-to-the-next-policy</guid>
      <pubDate>Tue, 06 Oct 2026 00:00:00 GMT</pubDate>
      <category>Benchmarks</category>
      <dc:creator>Sankalp Nagaonkar, Samuel Alexander</dc:creator>
      <description>A head-to-head retrieval benchmark, and what it means for robotics teams choosing what their next policy should learn from.</description>
    </item>
    <item>
      <title>Why a search tool is not enough for an LLM working with video</title>
      <link>https://videodb.io/blog/why-a-search-tool-is-not-enough-for-an-llm-working-with-video</link>
      <guid isPermaLink="true">https://videodb.io/blog/why-a-search-tool-is-not-enough-for-an-llm-working-with-video</guid>
      <pubDate>Mon, 05 Oct 2026 00:00:00 GMT</pubDate>
      <category>Agents</category>
      <dc:creator>Aditi Priya</dc:creator>
      <description>A search tool lets an LLM find a moment in video. But when the LLM is asked how many, or only the ones that failed, it counts whatever the search returned. We show why that happens and what the LLM needs instead.</description>
    </item>
    <item>
      <title>A query engine for robot video</title>
      <link>https://videodb.io/blog/robot-runs</link>
      <guid isPermaLink="true">https://videodb.io/blog/robot-runs</guid>
      <pubDate>Fri, 18 Sep 2026 00:00:00 GMT</pubDate>
      <category>Essay</category>
      <dc:creator>Ashutosh Trivedi</dc:creator>
      <description>We built a query engine that sits over a lab's robot video, keeps every episode whole, and answers questions in plain English: find the failures, export them as a training manifest, and search the next model's runs the moment they land.</description>
    </item>
    <item>
      <title>Can a VLM annotate a robot run? Results on a revised WGO-Bench</title>
      <link>https://videodb.io/blog/wgo-bench-action-annotations</link>
      <guid isPermaLink="true">https://videodb.io/blog/wgo-bench-action-annotations</guid>
      <pubDate>Tue, 08 Sep 2026 00:00:00 GMT</pubDate>
      <category>Benchmarks</category>
      <dc:creator>Samuel Alexander, Sankalp Nagaonkar</dc:creator>
      <description>We re-annotated all 100 videos of WGO-Bench, human and robot, then measured how well a VLM pipeline segments and labels the actions in them. VideoDB reaches 28.93% semantic F1 at IoU 0.5, ahead of the public baseline at 19.92%, and the remaining errors show where the work goes next.</description>
    </item>
    <item>
      <title>Claude edited our launch video with VideoDB skills</title>
      <link>https://videodb.io/blog/claude-edited-our-launch-video</link>
      <guid isPermaLink="true">https://videodb.io/blog/claude-edited-our-launch-video</guid>
      <pubDate>Sun, 09 Aug 2026 00:00:00 GMT</pubDate>
      <category>Agents</category>
      <dc:creator>Ashutosh Trivedi</dc:creator>
      <description>A case study in understanding-driven editing: the agent indexed the raw takes, picked the best take per narrative beat, composed and rendered the timeline, then re-indexed its own cut to audit it.</description>
    </item>
    <item>
      <title>Giving an agent a YouTube video it can search</title>
      <link>https://videodb.io/blog/ai-agent-watch-youtube</link>
      <guid isPermaLink="true">https://videodb.io/blog/ai-agent-watch-youtube</guid>
      <pubDate>Sun, 09 Aug 2026 00:00:00 GMT</pubDate>
      <category>Agents</category>
      <dc:creator>Ashutosh Trivedi</dc:creator>
      <description>Give your agent YouTube videos as searchable, timestamped context. Speech and visuals indexed, evidence clips compiled, in about 20 lines of Python.</description>
    </item>
    <item>
      <title>Twelve Labs alternatives, compared</title>
      <link>https://videodb.io/blog/twelve-labs-alternatives</link>
      <guid isPermaLink="true">https://videodb.io/blog/twelve-labs-alternatives</guid>
      <pubDate>Sun, 09 Aug 2026 00:00:00 GMT</pubDate>
      <category>Engineering</category>
      <dc:creator>Ashutosh Trivedi</dc:creator>
      <description>Comparing Twelve Labs alternatives for video AI in 2026: VideoDB, Mixpeek, Google, AWS, Azure, NVIDIA: organized by what you are actually building.</description>
    </item>
    <item>
      <title>Video RAG: how it works</title>
      <link>https://videodb.io/blog/video-rag</link>
      <guid isPermaLink="true">https://videodb.io/blog/video-rag</guid>
      <pubDate>Sun, 09 Aug 2026 00:00:00 GMT</pubDate>
      <category>Engineering</category>
      <dc:creator>Ashutosh Trivedi</dc:creator>
      <description>What video RAG is, why it breaks text-RAG assumptions, the four architectures that work, and how to build one, including RAG over live streams.</description>
    </item>
    <item>
      <title>Agentic Streams: an agent researches a topic and streams you the briefing</title>
      <link>https://videodb.io/blog/agentic-streams</link>
      <guid isPermaLink="true">https://videodb.io/blog/agentic-streams</guid>
      <pubDate>Sat, 08 Aug 2026 00:00:00 GMT</pubDate>
      <category>Agents</category>
      <dc:creator>Ashutosh Trivedi</dc:creator>
      <description>Open-source agents that research a topic on the live web, gather real clips, screenshots and charts, write a narration, assemble the video with VideoDB and hand back a stream you can play. Three agents included, with example outputs.</description>
    </item>
    <item>
      <title>Search over the Visual World: persistent visual memory, layered indexes, and source-grounded evidence</title>
      <link>https://videodb.io/blog/search-over-the-visual-world</link>
      <guid isPermaLink="true">https://videodb.io/blog/search-over-the-visual-world</guid>
      <pubDate>Sat, 08 Aug 2026 00:00:00 GMT</pubDate>
      <category>Papers</category>
      <dc:creator>Nagaonkar, Garg, Raj, Choithani, Trivedi</dc:creator>
      <description>VideoDB technical report (arXiv 2608.08075). Analyzer-defined scenes, persistent visual memory, capability-declared indexes, source-grounded evidence, and a 9,834-query retrieval evaluation against a commercial video-native engine.</description>
    </item>
    <item>
      <title>Real-time visual perception for agents</title>
      <link>https://videodb.io/blog/give-your-ai-agents-eyes</link>
      <guid isPermaLink="true">https://videodb.io/blog/give-your-ai-agents-eyes</guid>
      <pubDate>Fri, 07 Aug 2026 00:00:00 GMT</pubDate>
      <category>Agents</category>
      <dc:creator>Ashutosh Trivedi</dc:creator>
      <description>How AI agents get vision: ingest YouTube videos, live RTSP cameras, and screens; index continuously; search semantically; act on plain-language events.</description>
    </item>
    <item>
      <title>From an RTSP stream to events an agent can act on</title>
      <link>https://videodb.io/blog/rtsp-ai-analysis</link>
      <guid isPermaLink="true">https://videodb.io/blog/rtsp-ai-analysis</guid>
      <pubDate>Fri, 07 Aug 2026 00:00:00 GMT</pubDate>
      <category>Agents</category>
      <dc:creator>Ashutosh Trivedi</dc:creator>
      <description>Turn any RTSP camera stream into AI-readable events: continuous understanding, plain-language alerts, and searchable history. No CV pipeline to build.</description>
    </item>
    <item>
      <title>What video retrieval benchmarks taught us about ground truth</title>
      <link>https://videodb.io/blog/video-benchmark-ground-truth-ambiguity</link>
      <guid isPermaLink="true">https://videodb.io/blog/video-benchmark-ground-truth-ambiguity</guid>
      <pubDate>Fri, 24 Jul 2026 00:00:00 GMT</pubDate>
      <category>Benchmarks</category>
      <dc:creator>Samuel Alexander</dc:creator>
      <description>Manual review of MSRVTT, MSVD, VATEX, DiDeMo, and QVHighlights shows cases where the benchmark ground truth is too narrow, shifted, or ambiguous, so valid retrieved clips are scored as misses.</description>
    </item>
    <item>
      <title>DeepSeek’s visual primitives and the missing layer</title>
      <link>https://videodb.io/blog/deepseek-visual-primitives-reference-gap</link>
      <guid isPermaLink="true">https://videodb.io/blog/deepseek-visual-primitives-reference-gap</guid>
      <pubDate>Mon, 13 Jul 2026 00:00:00 GMT</pubDate>
      <category>Essay</category>
      <dc:creator>Sankalp Nagaonkar, Ashutosh Trivedi</dc:creator>
      <description>The Reference Gap explains why seeing an object is insufficient when a multimodal model cannot preserve its identity, location, or path through a long reasoning trace.</description>
    </item>
    <item>
      <title>JEPA, from language models to world models</title>
      <link>https://videodb.io/blog/jepa-from-language-models-to-world-models</link>
      <guid isPermaLink="true">https://videodb.io/blog/jepa-from-language-models-to-world-models</guid>
      <pubDate>Tue, 07 Jul 2026 00:00:00 GMT</pubDate>
      <category>Essay</category>
      <dc:creator>Sankalp Nagaonkar, Ashutosh Trivedi</dc:creator>
      <description>Why JEPA’s latent-prediction objective may shift AI systems from token prediction toward predictive world models for VLMs, VLAs, and embodied agents.</description>
    </item>
    <item>
      <title>TinyFish × VideoDB: the web your agent browses, as video it can search</title>
      <link>https://videodb.io/blog/tinyfish</link>
      <guid isPermaLink="true">https://videodb.io/blog/tinyfish</guid>
      <pubDate>Mon, 29 Jun 2026 00:00:00 GMT</pubDate>
      <category>Agents</category>
      <dc:creator>Ashutosh Trivedi</dc:creator>
      <description>TinyFish opens the web for agents. VideoDB lets them see inside video. Together they cover both access and comprehension.</description>
    </item>
    <item>
      <title>Episodic memory for agents</title>
      <link>https://videodb.io/blog/episodic-memory-for-agents</link>
      <guid isPermaLink="true">https://videodb.io/blog/episodic-memory-for-agents</guid>
      <pubDate>Thu, 11 Jun 2026 00:00:00 GMT</pubDate>
      <category>Essay</category>
      <dc:creator>Ashutosh Trivedi</dc:creator>
      <description>Humans remember experiences, not just facts. Your agent should too, with verifiable, playable evidence.</description>
    </item>
    <item>
      <title>Video was built for playback, not perception</title>
      <link>https://videodb.io/blog/playback-vs-perception</link>
      <guid isPermaLink="true">https://videodb.io/blog/playback-vs-perception</guid>
      <pubDate>Tue, 09 Jun 2026 00:00:00 GMT</pubDate>
      <category>Essay</category>
      <dc:creator>Ashutosh Trivedi</dc:creator>
      <description>70 years of video infrastructure for human eyes, and why AI needs perception-first architecture.</description>
    </item>
    <item>
      <title>MP4 is the wrong primitive</title>
      <link>https://videodb.io/blog/mp4-is-wrong-primitive</link>
      <guid isPermaLink="true">https://videodb.io/blog/mp4-is-wrong-primitive</guid>
      <pubDate>Sun, 07 Jun 2026 00:00:00 GMT</pubDate>
      <category>Essay</category>
      <dc:creator>Ashutosh Trivedi</dc:creator>
      <description>Video files were designed for playback. AI agents need indexes, not opaque blobs.</description>
    </item>
    <item>
      <title>Perception is the missing layer</title>
      <link>https://videodb.io/blog/perception-is-the-missing-layer</link>
      <guid isPermaLink="true">https://videodb.io/blog/perception-is-the-missing-layer</guid>
      <pubDate>Fri, 05 Jun 2026 00:00:00 GMT</pubDate>
      <category>Essay</category>
      <dc:creator>Ashutosh Trivedi</dc:creator>
      <description>LLMs have reasoning. RAG has retrieval. What's missing? The ability to perceive the world as it happens.</description>
    </item>
    <item>
      <title>Bloom: record your screen, search it later</title>
      <link>https://videodb.io/blog/bloom</link>
      <guid isPermaLink="true">https://videodb.io/blog/bloom</guid>
      <pubDate>Thu, 04 Jun 2026 00:00:00 GMT</pubDate>
      <category>Agents</category>
      <dc:creator>Lalit Gupta</dc:creator>
      <description>How I built a local-first screen recorder on top of the VideoDB SDK. It automatically indexes your audio and screen, making everything searchable by the time you paste the share link.</description>
    </item>
    <item>
      <title>Focusd: a coach for your workday</title>
      <link>https://videodb.io/blog/focusd</link>
      <guid isPermaLink="true">https://videodb.io/blog/focusd</guid>
      <pubDate>Thu, 04 Jun 2026 00:00:00 GMT</pubDate>
      <category>Engineering</category>
      <dc:creator>Sankalp Nagaonkar</dc:creator>
      <description>How we built Focusd as a local-first desktop app that records work sessions, indexes screen activity with VideoDB, and turns raw events into useful productivity coaching.</description>
    </item>
    <item>
      <title>A camera feed for your OpenClaw agent</title>
      <link>https://videodb.io/blog/openclaw-agent-camera</link>
      <guid isPermaLink="true">https://videodb.io/blog/openclaw-agent-camera</guid>
      <pubDate>Thu, 04 Jun 2026 00:00:00 GMT</pubDate>
      <category>Agents</category>
      <dc:creator>Ashutosh Trivedi</dc:creator>
      <description>How we built a VideoDB-powered OpenClaw skill that records, indexes, searches, summarizes, and clips an agent's remote desktop without changing the agent itself.</description>
    </item>
    <item>
      <title>Infrastructure that sees and edits</title>
      <link>https://videodb.io/blog/infrastructure-that-sees-and-edits</link>
      <guid isPermaLink="true">https://videodb.io/blog/infrastructure-that-sees-and-edits</guid>
      <pubDate>Wed, 03 Jun 2026 00:00:00 GMT</pubDate>
      <category>Essay</category>
      <dc:creator>Ashutosh Trivedi</dc:creator>
      <description>Explore VideoDB's multimodal infrastructure that combines vision and editing for intelligent video processing.</description>
    </item>
    <item>
      <title>Why agents are blind today</title>
      <link>https://videodb.io/blog/why-agents-are-blind</link>
      <guid isPermaLink="true">https://videodb.io/blog/why-agents-are-blind</guid>
      <pubDate>Mon, 01 Jun 2026 00:00:00 GMT</pubDate>
      <category>Essay</category>
      <dc:creator>Ashutosh Trivedi</dc:creator>
      <description>The gap between human perception and agent perception, and why it matters for the future of AI.</description>
    </item>
    <item>
      <title>A strong general-purpose VLM still fails on a chessboard</title>
      <link>https://videodb.io/blog/claude-chessboard-spatial-reasoning</link>
      <guid isPermaLink="true">https://videodb.io/blog/claude-chessboard-spatial-reasoning</guid>
      <pubDate>Tue, 19 May 2026 00:00:00 GMT</pubDate>
      <category>Benchmarks</category>
      <dc:creator>Sankalp Nagaonkar</dc:creator>
      <description>We tested GPT-5.4 and Claude Opus 4.7 on reading 72 chess positions into FEN. Claude Opus 4.7 understood the boards and still lost to square-level errors, and the reason only showed up in the reasoning traces.</description>
    </item>
    <item>
      <title>Deep Search: finding exact moments in video</title>
      <link>https://videodb.io/blog/deepsearch</link>
      <guid isPermaLink="true">https://videodb.io/blog/deepsearch</guid>
      <pubDate>Tue, 19 May 2026 00:00:00 GMT</pubDate>
      <category>Agents</category>
      <dc:creator>Sankalp Nagaonkar</dc:creator>
      <description>How we built Deep Search as a retrieval loop for finding exact moments in video using planning, indexing, validation, recovery, and follow-up state.</description>
    </item>
    <item>
      <title>Pair Programmer: live screen and audio context for coding agents</title>
      <link>https://videodb.io/blog/pair-programmer-live-context</link>
      <guid isPermaLink="true">https://videodb.io/blog/pair-programmer-live-context</guid>
      <pubDate>Tue, 19 May 2026 00:00:00 GMT</pubDate>
      <category>Agents</category>
      <dc:creator>Rohit Garg</dc:creator>
      <description>A skill that gives Claude Code, Cursor and Codex live screen, mic and system-audio context, and how it is built on VideoDB Capture, RTStreams, local events and agent-side retrieval.</description>
    </item>
    <item>
      <title>Evaluating video VLMs on your own task</title>
      <link>https://videodb.io/blog/how-to-evaluate-multimodal-vlms-for-your-video-use-case</link>
      <guid isPermaLink="true">https://videodb.io/blog/how-to-evaluate-multimodal-vlms-for-your-video-use-case</guid>
      <pubDate>Fri, 15 May 2026 00:00:00 GMT</pubDate>
      <category>Benchmarks</category>
      <dc:creator>Sankalp Nagaonkar</dc:creator>
      <description>A practical workflow for evaluating video VLM setups with VideoDB and Langfuse, from task definition and dataset design to tracing, scoring, and deployment decisions.</description>
    </item>
    <item>
      <title>Dispatch #001: playable evidence, thinking tokens, agentic streams</title>
      <link>https://videodb.io/blog/dispatch-001</link>
      <guid isPermaLink="true">https://videodb.io/blog/dispatch-001</guid>
      <pubDate>Fri, 08 May 2026 00:00:00 GMT</pubDate>
      <category>Engineering</category>
      <dc:creator>Ashutosh Trivedi</dc:creator>
      <description>The first issue of the VideoDB Dispatch: why agent runs need playable evidence, how much a video model should think, and builder updates from VideoDB.</description>
    </item>
    <item>
      <title>How I built Call.md on VideoDB</title>
      <link>https://videodb.io/blog/how-i-built-call-intelligence-with-videodb</link>
      <guid isPermaLink="true">https://videodb.io/blog/how-i-built-call-intelligence-with-videodb</guid>
      <pubDate>Fri, 08 May 2026 00:00:00 GMT</pubDate>
      <category>Agents</category>
      <dc:creator>Om</dc:creator>
      <description>A practical build story for a local-first call intelligence app that records calls, transcribes speakers, generates live nudges, and exports structured Markdown.</description>
    </item>
    <item>
      <title>Reuse HTTP connections in Python services</title>
      <link>https://videodb.io/blog/python-http-connection-pooling</link>
      <guid isPermaLink="true">https://videodb.io/blog/python-http-connection-pooling</guid>
      <pubDate>Mon, 20 Apr 2026 00:00:00 GMT</pubDate>
      <category>Engineering</category>
      <dc:creator>Rohit Garg</dc:creator>
      <description>A short field note on replacing repeated bare requests calls with a shared session so Python services can reuse HTTP connections.</description>
    </item>
    <item>
      <title>Caching CORS preflights: hundreds of milliseconds per call</title>
      <link>https://videodb.io/blog/cors-preflight-cache</link>
      <guid isPermaLink="true">https://videodb.io/blog/cors-preflight-cache</guid>
      <pubDate>Fri, 17 Apr 2026 00:00:00 GMT</pubDate>
      <category>Engineering</category>
      <dc:creator>Om Gate</dc:creator>
      <description>A practical note on preflight caching, custom auth headers, and why OPTIONS requests quietly dominate app latency on slower connections.</description>
    </item>
    <item>
      <title>A Postgres backup is not real until you restore it</title>
      <link>https://videodb.io/blog/postgres-backup-restore-drill</link>
      <guid isPermaLink="true">https://videodb.io/blog/postgres-backup-restore-drill</guid>
      <pubDate>Fri, 02 Jan 2026 00:00:00 GMT</pubDate>
      <category>Engineering</category>
      <dc:creator>Rohit Garg</dc:creator>
      <description>A field note from a simple offline Postgres migration: why pg_dump was enough, where read-only mode fits, and why the restore drill matters more than the dump command.</description>
    </item>
    <item>
      <title>The 6 MB Lambda limit: compress before, not after</title>
      <link>https://videodb.io/blog/lambda-compression-trap</link>
      <guid isPermaLink="true">https://videodb.io/blog/lambda-compression-trap</guid>
      <pubDate>Wed, 17 Sep 2025 00:00:00 GMT</pubDate>
      <category>Engineering</category>
      <dc:creator>Rohit Garg</dc:creator>
      <description>Why API Gateway compression can make large responses look safe in testing, then still fail when Lambda enforces the raw payload limit first.</description>
    </item>
    <item>
      <title>Searching the slides in a conference recording</title>
      <link>https://videodb.io/blog/conference-slide-extraction</link>
      <guid isPermaLink="true">https://videodb.io/blog/conference-slide-extraction</guid>
      <pubDate>Mon, 25 Aug 2025 00:00:00 GMT</pubDate>
      <category>Engineering</category>
      <dc:creator>Ashutosh Trivedi</dc:creator>
      <description>Pull slide content out of any conference talk by combining spoken-word search with visual scene indexing. Full Python pipeline with working code.</description>
    </item>
    <item>
      <title>NFL game analysis: 80% fewer VLM hallucinations</title>
      <link>https://videodb.io/blog/nfl-game-analysis-vlm-hallucinations</link>
      <guid isPermaLink="true">https://videodb.io/blog/nfl-game-analysis-vlm-hallucinations</guid>
      <pubDate>Mon, 25 Aug 2025 00:00:00 GMT</pubDate>
      <category>Benchmarks</category>
      <dc:creator>Sankalp Nagaonkar</dc:creator>
      <description>Three approaches to event-dense NFL footage, measured. Play-by-play segmentation cut VLM hallucinations from 68.1% to 11.4% at up to 70% lower cost.</description>
    </item>
    <item>
      <title>Smaller Python ML containers</title>
      <link>https://videodb.io/blog/python-ml-container-size</link>
      <guid isPermaLink="true">https://videodb.io/blog/python-ml-container-size</guid>
      <pubDate>Mon, 25 Aug 2025 00:00:00 GMT</pubDate>
      <category>Engineering</category>
      <dc:creator>Lalit Gupta</dc:creator>
      <description>Slim bases, same-layer cleanup, CPU-only PyTorch, pinned wheels, and the unglamorous work of keeping deploys fast.</description>
    </item>
    <item>
      <title>VideoDB × TwelveLabs: search across your video library</title>
      <link>https://videodb.io/blog/twelvelabs</link>
      <guid isPermaLink="true">https://videodb.io/blog/twelvelabs</guid>
      <pubDate>Wed, 13 Aug 2025 00:00:00 GMT</pubDate>
      <category>Engineering</category>
      <dc:creator>Ashutosh Trivedi</dc:creator>
      <description>TwelveLabs multimodal understanding plugs into VideoDB so any agent can find the exact moment it needs, not just the right file.</description>
    </item>
    <item>
      <title>VideoDB × LlamaIndex: video in your RAG pipeline</title>
      <link>https://videodb.io/blog/llama-index</link>
      <guid isPermaLink="true">https://videodb.io/blog/llama-index</guid>
      <pubDate>Wed, 11 Jun 2025 00:00:00 GMT</pubDate>
      <category>Engineering</category>
      <dc:creator>Ashutosh Trivedi</dc:creator>
      <description>The official VideoDB connector for LlamaIndex lets you treat video as a first-class source in any retrieval-augmented pipeline.</description>
    </item>
  </channel>
</rss>