[HN Gopher] Show HN: Gemini can now natively embed video, so I b...
       ___________________________________________________________________
        
       Show HN: Gemini can now natively embed video, so I built sub-second
       video search
        
       Gemini Embedding 2 can project raw video directly into a
       768-dimensional vector space alongside text. No transcription, no
       frame captioning, no intermediate text. A query like "green car
       cutting me off" is directly comparable to a 30-second video clip at
       the vector level.  I used this to build a CLI that indexes hours of
       footage into ChromaDB, then searches it with natural language and
       auto-trims the matching clip. Demo video on the GitHub README.
       Indexing costs ~$2.50/hr of footage. Still-frame detection skips
       idle chunks, so security camera / sentry mode footage is much
       cheaper.
        
       Author : sohamrj
       Score  : 207 points
       Date   : 2026-03-24 14:58 UTC (8 hours ago)
        
 (HTM) web link (github.com)
 (TXT) w3m dump (github.com)
        
       | ygouzerh wrote:
       | That's quite interesting, well done! I haven't thought of this
       | use case for embeddings. It open the door to quite many potential
       | applications!
        
         | stavros wrote:
         | Man, the surveillance applications for this are staggering.
        
       | dev_tools_lab wrote:
       | Nice use of native video embedding. How do you handle cases where
       | Gemini's response confidence is low? Do you have a fallback or
       | threshold?
        
         | sohamrj wrote:
         | as of now, no threshold but that is planned in the future.
         | 
         | for example, for now if i search "cybertruck" in my indexed
         | dashcam footage, i don't have any cybertrucks in my footage, so
         | it'll return a clip of the next best match which is a big
         | truck, but not a cybertruck
        
       | mdrzn wrote:
       | Very interesting (not for a dashcam, but for home monitoring).
        
       | klntsky wrote:
       | why not skip the text conversion? is it usable at all?
        
         | sohamrj wrote:
         | gemini embedding 2 converts straight video to vectors. in this
         | case, dashcam clips don't have audio to transcribe and even if
         | they did, it would be useless in the search
        
           | password4321 wrote:
           | What are the SoA audio models right now?
        
       | Aeroi wrote:
       | very cool, anybody have apparent use cases for this?
        
         | sohamrj wrote:
         | dashcam and home security footage are the 2 main ones i can
         | think of.
         | 
         | a bit expensive right now so it's not as practical at scale.
         | but once the embedding model comes out of public preview, and
         | we hopefully get a local equivalent, this will be a lot more
         | practical.
        
         | hebelehubele wrote:
         | State surveillance
        
           | wahnfrieden wrote:
           | Worker surveillance
        
         | giozaarour wrote:
         | I think a good use case would be searching for certain products
         | or videos across social media (TikTok and Instagram).
         | especially useful for shopping, maybe
        
           | vidarh wrote:
           | Branding/marketing monitoring companies would be all over
           | this.
        
       | emsign wrote:
       | Where is the Exit to this dystopia?
        
         | BrokenCogs wrote:
         | The Matrix style human pods: we live in blissful ignorance in
         | the Matrix, while the LLMs extract more and more compute power
         | from us so some CEO somewhere can claim they have now replaced
         | all humans with machines in their business.
        
           | throwup238 wrote:
           | I was thinking more of the season 3 episode of Doctor Who
           | titled _Gridlock_ where everyone lives in flying cars
           | circling a giant expressway underground, while all the upper
           | class people on the surface died years ago from a pandemic.
        
           | ting0 wrote:
           | Ever get the feeling that the universe is reading your mind?
           | Maybe there's some truth to that after all.
        
         | RobotToaster wrote:
         | In the matrix the exit was pay phones, which perhaps explains
         | why our overlords are removing them
        
         | jama211 wrote:
         | I don't think this means we're in a dystopia
        
           | zwirbl wrote:
           | You might not have been paying attention
        
             | 52-6F-62 wrote:
             | I think Radiohead said that
        
         | draw_down wrote:
         | The dystopia of searching for video clips and finding them?
         | What?
        
           | bitexploder wrote:
           | Yes? Right now it is relatively expensive to search video. As
           | embedding tech like this advances and makes it even cheaper
           | it just increases the ability to search and analyze every
           | movement. "Locate speech patterns that indicate dissident
           | activity using the dissident activity skill"
        
         | nclin_ wrote:
         | Well, with data analysis powers like this a few treasonous
         | words in front of a flock camera will show you the way.
        
         | anxoo wrote:
         | https://pauseai.info/
        
       | 7777777phil wrote:
       | Today I learned that Gemini can now natively embed video..
       | 
       | Cool Project, thanks for sharing!
        
       | kamranjon wrote:
       | Does anyone know of an open weights models that can embed video?
       | Would love to experiment locally with this.
        
         | sohamrj wrote:
         | Not aware of any that do native video-to-vector embedding the
         | way Gemini Embedding 2 does. There are CLIP-based models (like
         | VideoCLIP) that embed frames individually, but they don't
         | process temporal video. you'd need to average frame embeddings
         | which loses a lot.
         | 
         | Would love to see open-weight models with this capability since
         | it would eliminate the API cost and the privacy concern of
         | uploading footage.
        
       | SpaceManNabs wrote:
       | > No transcription, no frame captioning, no intermediate text.
       | 
       | If there is text on the video (like a caption or wtv), will the
       | embedding capture that? Never thought about this before.
       | 
       | If the video has audio, does the embedding capture that too?
        
         | sohamrj wrote:
         | Yes to both. The embedding is over raw video frames, so
         | anything visible (text, signs, captions) gets captured in the
         | vector. And Gemini Embedding 2 extracts the audio track and
         | embeds it alongside the visual frames. So a query like 'someone
         | yelling' would theoretically match on audio. My dashcam footage
         | doesn't have audio though, so I haven't tested that side yet.
        
       | nullbyte wrote:
       | What a brilliant idea! is this all done locally? That's
       | incredible.
        
         | apwheele wrote:
         | While the vector store is local, it is sending the data to
         | Gemini's API for embedding. (Which if using a paid API key is
         | probably fine for most use cases, no long term
         | retention/training etc.)
        
       | simonreiff wrote:
       | Very impressive! A webhook could be configured to trigger an
       | alarm if a semantic match to any category of activities is
       | detected, and then you basically have a virtual security guard
       | and private investigator. Well played.
        
         | sohamrj wrote:
         | Thanks! Yeah that would be pretty cool, but continuous indexing
         | would be pretty expensive now, because the model's in public
         | preview and there are no local alternatives afaik.
         | 
         | This very well might be a reality in a couple years though!
        
       | macNchz wrote:
       | This is a really cool implementation--embeddings still often feel
       | like magic to me. That said, this exact use case is sort of also
       | my biggest point of concern with where AI takes us, much more so
       | than most of the common AI risks you hear lots of chatter about.
       | We live in a world absolutely loaded with cameras now but
       | ultimately retain some semblance of semi-anonymity/privacy in
       | public by virtue of the fact that nobody can actually watch or
       | review all of the video from those cameras except when there is a
       | compelling reason to do so, but these technologies are making
       | that a much more realistic proposition.
       | 
       | The presence of cameras everywhere is considerably more
       | concerning than the status quo, to me at least, when there is an
       | AI watching and indexing every second of every feed--where camera
       | owners or manufacturers or governments could set simple natural
       | language parameters for highly specific people or activities
       | notify about. There are obviously compelling and easy-to-sell
       | cases here that will surely drive adoption as it becomes cost
       | effective: get an alert to crime in progress, get an alert when a
       | neighbor who doesn't clean up after his dog, get an alert when
       | someone has fallen...but the potential implications of living in
       | a panopticon like this if not well regulated are pretty ugly.
        
         | sohamrj wrote:
         | Totally valid concern. Right now the cost ($2.50/hr) and
         | latency make continuous real-time indexing impractical, but
         | that won't always be the case. This is one of the reasons I'd
         | want to see open-weight local models for this, keeps the
         | indexing on your own hardware with no footage leaving your
         | machine. But you're right that the broader trajectory here is
         | worth thinking carefully about.
        
           | mpalmer wrote:
           | It's 2.50 an hour because Google has margins. A nation state
           | could do it at cost, and even if it's not a huge difference,
           | the price of a year's worth of embeddings is just $21,900.
           | That's a rounding error, especially considering it's a one
           | time cost for footage.
        
             | wholinator2 wrote:
             | Right? $2.50 an hour is trivial to a Government that can
             | vote to invent a trillion dollars. Even just 1 million
             | dollars is the cost of monitoring 45 real time feeds for a
             | year. I'm sure just many very rich people would pay that
             | for the safety of their compound.
        
           | jimmySixDOF wrote:
           | How are you getting to $2.50/hr ? The price sheet says its
           | 0.00079 per frame.
           | 
           | https://ai.google.dev/gemini-api/docs/pricing#gemini-
           | embeddi...
        
         | citruscomputing wrote:
         | It's being built as we speak. I attended at a city council
         | meeting yesterday, discussing approving a contract for ALPR
         | cameras. I learned about a product from the camera vendor
         | called Fusus[0], a dashboard that integrates various camera
         | systems, ALPRs, alerts, etc. Two things stood out to me:
         | natural-language querying of video feeds, and future planned
         | integration with civilian-deployed cameras. The city only had
         | budget for 50 ALPRs, and they stressed how they're only
         | deploying them on main streets, but it seems like only a matter
         | of time before your neighbor is able to install a camera that
         | feeds right into the local PD's AI-enabled systems. One council
         | member raised concerns about integrations with the citizen
         | app[1] specifically (and a few others I didn't catch the names
         | of). I'm very worried about where all this is heading.
         | 
         | [0]: https://www.axon.com/products/axon-fusus [1]:
         | https://citizen.com/
        
         | cake_robot wrote:
         | Yeah, the panopticon is now technically very feasible it's just
         | expensive to implement (for now).
        
         | Ajedi32 wrote:
         | Most cameras are also not queryable by any one person or
         | organization. They are owned by different companies and if the
         | government wants access they have to subpoena them after the
         | fact.
         | 
         | The problems start cropping up when you get things like Flock
         | where governments start deploying cameras on a massive scale,
         | or Ring where a single company has unrestricted access to
         | everyone's private cameras.
        
           | Spivak wrote:
           | I think Flock is just a symptom of the underlying tech
           | becoming so cheap that "just blanket the city in cameras"
           | starts to sound like a viable solution when police rely so
           | heavily on camera footage.
           | 
           | I don't think it's a good thing but it seems the limiting
           | factor has been technological feasibility instead of any kind
           | of principle against it.
        
         | janalsncm wrote:
         | For specific people they probably wouldn't use general
         | embeddings. These embeddings can let you search for "tall man
         | in a trenchcoat" but if you want a specific person you would
         | use facial recognition.
        
           | hypeatei wrote:
           | I think a general description is better for
           | surveillance/tracking like this, no? If they're at a weird
           | angle or intentionally concealing their face then facial
           | recognition falls apart but being able to describe them
           | naturally would result in better tracking IMO.
        
             | macNchz wrote:
             | Presumably the ideal is some kind of a fusion. Upload or
             | tag some images/videos and link someone's social profiles
             | and the system can look out for them based on facial
             | recognition, gait recognition, vehicle/pets/common wardrobe
             | items in combination.
        
         | greggsy wrote:
         | All the major cloud providers offer some form of face detection
         | and numberplate reading, with many supporting object detection
         | (ie package, vehicle, person) out of the camera itself.
        
           | macNchz wrote:
           | It's definitely creeping into things, though most of the
           | features I've seen are fairly simplistic compared to what
           | would be possible if the video was being reviewed + indexed
           | by current SoTA multimodal LLMs.
        
       | danbrooks wrote:
       | I work in content/video intelligence. Gemini is great for this
       | type of use case out of the box.
        
       | cloogshicer wrote:
       | Could this be used for creating video editing software?
       | 
       | Imagine a Premiere plugin where you could say "remove all scenes
       | containing cats" and it'll spit out an EDL (Edit Decision List)
       | that you can still manually adjust.
        
       | rigrassm wrote:
       | I picked up a Rexing dash cam a few months back and after getting
       | frustrated with how clunky it is to get footage of it, I decided
       | to look into building something out myself to browse and download
       | the recordings without having to pull the SD card. While
       | scrolling through the recordings, I explicitly remember thinking
       | it would be nice to just describe what I was looking for and run
       | a search. Looking forward to incorporating this into my project.
       | 
       | Thanks for sharing!
        
       | totisjosema wrote:
       | What is your experience so far with the quality of the retrieved
       | pieces?
        
         | sohamrj wrote:
         | I've found I have to be very specific to get the clip I'm
         | searching for. For example, "car cuts me off" just returned a
         | clip of a car driving past my blindspot. But, "car with bike
         | rack on back cuts me off at night" gave me exactly the clip I
         | was looking for.
        
       | bobafett-9902 wrote:
       | I wonder if the underlying improvements in visual language
       | learning will allow for even more efficient search. The First
       | Fully General Computer Action Model -> https://si.inc/posts/fdm1/
        
       | QubridAI wrote:
       | This is a big leap true multimodal search without text
       | bottlenecks makes video querying feel finally native and insanely
       | practical.
        
       | WatchDog wrote:
       | I don't quite understand the 5 second overlap. I assume it's so
       | that events that occur over the chunk boundary don't get missed,
       | but is there any examples or benchmarking to examine how useful
       | this is?
        
         | sohamrj wrote:
         | yea, it's so events on a chunk boundary still get captured in
         | at least one chunk. i haven't had the chance to do formal
         | benchmarks on overlap vs. no-overlap yet. the 5s default is a
         | pragmatic choice, long enough to catch most events that would
         | otherwise be split, short enough to not add much cost (120
         | chunks/hr to ~138). also it's configurable via the --overlap
         | flag.
        
       | thegabriele wrote:
       | Why just the dash cam?
        
         | sohamrj wrote:
         | dashcam is just one of the use cases and the one i tested on.
         | but this could theoretically work with any kind of video
         | footage like home security footage
        
       | lwarfield wrote:
       | Damn, I need to going with my embeddings project. I've currently
       | got a prototype for using embeddings (not gemini in my case) for
       | making a game that's kinda reverse connections:
       | 
       | collections.lwarfield.dev
        
       ___________________________________________________________________
       (page generated 2026-03-24 23:00 UTC)