[HN Gopher] Show HN: Gemini can now natively embed video, so I b...
___________________________________________________________________
Show HN: Gemini can now natively embed video, so I built sub-second
video search
Gemini Embedding 2 can project raw video directly into a
768-dimensional vector space alongside text. No transcription, no
frame captioning, no intermediate text. A query like "green car
cutting me off" is directly comparable to a 30-second video clip at
the vector level. I used this to build a CLI that indexes hours of
footage into ChromaDB, then searches it with natural language and
auto-trims the matching clip. Demo video on the GitHub README.
Indexing costs ~$2.50/hr of footage. Still-frame detection skips
idle chunks, so security camera / sentry mode footage is much
cheaper.
Author : sohamrj
Score : 207 points
Date : 2026-03-24 14:58 UTC (8 hours ago)
(HTM) web link (github.com)
(TXT) w3m dump (github.com)
| ygouzerh wrote:
| That's quite interesting, well done! I haven't thought of this
| use case for embeddings. It open the door to quite many potential
| applications!
| stavros wrote:
| Man, the surveillance applications for this are staggering.
| dev_tools_lab wrote:
| Nice use of native video embedding. How do you handle cases where
| Gemini's response confidence is low? Do you have a fallback or
| threshold?
| sohamrj wrote:
| as of now, no threshold but that is planned in the future.
|
| for example, for now if i search "cybertruck" in my indexed
| dashcam footage, i don't have any cybertrucks in my footage, so
| it'll return a clip of the next best match which is a big
| truck, but not a cybertruck
| mdrzn wrote:
| Very interesting (not for a dashcam, but for home monitoring).
| klntsky wrote:
| why not skip the text conversion? is it usable at all?
| sohamrj wrote:
| gemini embedding 2 converts straight video to vectors. in this
| case, dashcam clips don't have audio to transcribe and even if
| they did, it would be useless in the search
| password4321 wrote:
| What are the SoA audio models right now?
| Aeroi wrote:
| very cool, anybody have apparent use cases for this?
| sohamrj wrote:
| dashcam and home security footage are the 2 main ones i can
| think of.
|
| a bit expensive right now so it's not as practical at scale.
| but once the embedding model comes out of public preview, and
| we hopefully get a local equivalent, this will be a lot more
| practical.
| hebelehubele wrote:
| State surveillance
| wahnfrieden wrote:
| Worker surveillance
| giozaarour wrote:
| I think a good use case would be searching for certain products
| or videos across social media (TikTok and Instagram).
| especially useful for shopping, maybe
| vidarh wrote:
| Branding/marketing monitoring companies would be all over
| this.
| emsign wrote:
| Where is the Exit to this dystopia?
| BrokenCogs wrote:
| The Matrix style human pods: we live in blissful ignorance in
| the Matrix, while the LLMs extract more and more compute power
| from us so some CEO somewhere can claim they have now replaced
| all humans with machines in their business.
| throwup238 wrote:
| I was thinking more of the season 3 episode of Doctor Who
| titled _Gridlock_ where everyone lives in flying cars
| circling a giant expressway underground, while all the upper
| class people on the surface died years ago from a pandemic.
| ting0 wrote:
| Ever get the feeling that the universe is reading your mind?
| Maybe there's some truth to that after all.
| RobotToaster wrote:
| In the matrix the exit was pay phones, which perhaps explains
| why our overlords are removing them
| jama211 wrote:
| I don't think this means we're in a dystopia
| zwirbl wrote:
| You might not have been paying attention
| 52-6F-62 wrote:
| I think Radiohead said that
| draw_down wrote:
| The dystopia of searching for video clips and finding them?
| What?
| bitexploder wrote:
| Yes? Right now it is relatively expensive to search video. As
| embedding tech like this advances and makes it even cheaper
| it just increases the ability to search and analyze every
| movement. "Locate speech patterns that indicate dissident
| activity using the dissident activity skill"
| nclin_ wrote:
| Well, with data analysis powers like this a few treasonous
| words in front of a flock camera will show you the way.
| anxoo wrote:
| https://pauseai.info/
| 7777777phil wrote:
| Today I learned that Gemini can now natively embed video..
|
| Cool Project, thanks for sharing!
| kamranjon wrote:
| Does anyone know of an open weights models that can embed video?
| Would love to experiment locally with this.
| sohamrj wrote:
| Not aware of any that do native video-to-vector embedding the
| way Gemini Embedding 2 does. There are CLIP-based models (like
| VideoCLIP) that embed frames individually, but they don't
| process temporal video. you'd need to average frame embeddings
| which loses a lot.
|
| Would love to see open-weight models with this capability since
| it would eliminate the API cost and the privacy concern of
| uploading footage.
| SpaceManNabs wrote:
| > No transcription, no frame captioning, no intermediate text.
|
| If there is text on the video (like a caption or wtv), will the
| embedding capture that? Never thought about this before.
|
| If the video has audio, does the embedding capture that too?
| sohamrj wrote:
| Yes to both. The embedding is over raw video frames, so
| anything visible (text, signs, captions) gets captured in the
| vector. And Gemini Embedding 2 extracts the audio track and
| embeds it alongside the visual frames. So a query like 'someone
| yelling' would theoretically match on audio. My dashcam footage
| doesn't have audio though, so I haven't tested that side yet.
| nullbyte wrote:
| What a brilliant idea! is this all done locally? That's
| incredible.
| apwheele wrote:
| While the vector store is local, it is sending the data to
| Gemini's API for embedding. (Which if using a paid API key is
| probably fine for most use cases, no long term
| retention/training etc.)
| simonreiff wrote:
| Very impressive! A webhook could be configured to trigger an
| alarm if a semantic match to any category of activities is
| detected, and then you basically have a virtual security guard
| and private investigator. Well played.
| sohamrj wrote:
| Thanks! Yeah that would be pretty cool, but continuous indexing
| would be pretty expensive now, because the model's in public
| preview and there are no local alternatives afaik.
|
| This very well might be a reality in a couple years though!
| macNchz wrote:
| This is a really cool implementation--embeddings still often feel
| like magic to me. That said, this exact use case is sort of also
| my biggest point of concern with where AI takes us, much more so
| than most of the common AI risks you hear lots of chatter about.
| We live in a world absolutely loaded with cameras now but
| ultimately retain some semblance of semi-anonymity/privacy in
| public by virtue of the fact that nobody can actually watch or
| review all of the video from those cameras except when there is a
| compelling reason to do so, but these technologies are making
| that a much more realistic proposition.
|
| The presence of cameras everywhere is considerably more
| concerning than the status quo, to me at least, when there is an
| AI watching and indexing every second of every feed--where camera
| owners or manufacturers or governments could set simple natural
| language parameters for highly specific people or activities
| notify about. There are obviously compelling and easy-to-sell
| cases here that will surely drive adoption as it becomes cost
| effective: get an alert to crime in progress, get an alert when a
| neighbor who doesn't clean up after his dog, get an alert when
| someone has fallen...but the potential implications of living in
| a panopticon like this if not well regulated are pretty ugly.
| sohamrj wrote:
| Totally valid concern. Right now the cost ($2.50/hr) and
| latency make continuous real-time indexing impractical, but
| that won't always be the case. This is one of the reasons I'd
| want to see open-weight local models for this, keeps the
| indexing on your own hardware with no footage leaving your
| machine. But you're right that the broader trajectory here is
| worth thinking carefully about.
| mpalmer wrote:
| It's 2.50 an hour because Google has margins. A nation state
| could do it at cost, and even if it's not a huge difference,
| the price of a year's worth of embeddings is just $21,900.
| That's a rounding error, especially considering it's a one
| time cost for footage.
| wholinator2 wrote:
| Right? $2.50 an hour is trivial to a Government that can
| vote to invent a trillion dollars. Even just 1 million
| dollars is the cost of monitoring 45 real time feeds for a
| year. I'm sure just many very rich people would pay that
| for the safety of their compound.
| jimmySixDOF wrote:
| How are you getting to $2.50/hr ? The price sheet says its
| 0.00079 per frame.
|
| https://ai.google.dev/gemini-api/docs/pricing#gemini-
| embeddi...
| citruscomputing wrote:
| It's being built as we speak. I attended at a city council
| meeting yesterday, discussing approving a contract for ALPR
| cameras. I learned about a product from the camera vendor
| called Fusus[0], a dashboard that integrates various camera
| systems, ALPRs, alerts, etc. Two things stood out to me:
| natural-language querying of video feeds, and future planned
| integration with civilian-deployed cameras. The city only had
| budget for 50 ALPRs, and they stressed how they're only
| deploying them on main streets, but it seems like only a matter
| of time before your neighbor is able to install a camera that
| feeds right into the local PD's AI-enabled systems. One council
| member raised concerns about integrations with the citizen
| app[1] specifically (and a few others I didn't catch the names
| of). I'm very worried about where all this is heading.
|
| [0]: https://www.axon.com/products/axon-fusus [1]:
| https://citizen.com/
| cake_robot wrote:
| Yeah, the panopticon is now technically very feasible it's just
| expensive to implement (for now).
| Ajedi32 wrote:
| Most cameras are also not queryable by any one person or
| organization. They are owned by different companies and if the
| government wants access they have to subpoena them after the
| fact.
|
| The problems start cropping up when you get things like Flock
| where governments start deploying cameras on a massive scale,
| or Ring where a single company has unrestricted access to
| everyone's private cameras.
| Spivak wrote:
| I think Flock is just a symptom of the underlying tech
| becoming so cheap that "just blanket the city in cameras"
| starts to sound like a viable solution when police rely so
| heavily on camera footage.
|
| I don't think it's a good thing but it seems the limiting
| factor has been technological feasibility instead of any kind
| of principle against it.
| janalsncm wrote:
| For specific people they probably wouldn't use general
| embeddings. These embeddings can let you search for "tall man
| in a trenchcoat" but if you want a specific person you would
| use facial recognition.
| hypeatei wrote:
| I think a general description is better for
| surveillance/tracking like this, no? If they're at a weird
| angle or intentionally concealing their face then facial
| recognition falls apart but being able to describe them
| naturally would result in better tracking IMO.
| macNchz wrote:
| Presumably the ideal is some kind of a fusion. Upload or
| tag some images/videos and link someone's social profiles
| and the system can look out for them based on facial
| recognition, gait recognition, vehicle/pets/common wardrobe
| items in combination.
| greggsy wrote:
| All the major cloud providers offer some form of face detection
| and numberplate reading, with many supporting object detection
| (ie package, vehicle, person) out of the camera itself.
| macNchz wrote:
| It's definitely creeping into things, though most of the
| features I've seen are fairly simplistic compared to what
| would be possible if the video was being reviewed + indexed
| by current SoTA multimodal LLMs.
| danbrooks wrote:
| I work in content/video intelligence. Gemini is great for this
| type of use case out of the box.
| cloogshicer wrote:
| Could this be used for creating video editing software?
|
| Imagine a Premiere plugin where you could say "remove all scenes
| containing cats" and it'll spit out an EDL (Edit Decision List)
| that you can still manually adjust.
| rigrassm wrote:
| I picked up a Rexing dash cam a few months back and after getting
| frustrated with how clunky it is to get footage of it, I decided
| to look into building something out myself to browse and download
| the recordings without having to pull the SD card. While
| scrolling through the recordings, I explicitly remember thinking
| it would be nice to just describe what I was looking for and run
| a search. Looking forward to incorporating this into my project.
|
| Thanks for sharing!
| totisjosema wrote:
| What is your experience so far with the quality of the retrieved
| pieces?
| sohamrj wrote:
| I've found I have to be very specific to get the clip I'm
| searching for. For example, "car cuts me off" just returned a
| clip of a car driving past my blindspot. But, "car with bike
| rack on back cuts me off at night" gave me exactly the clip I
| was looking for.
| bobafett-9902 wrote:
| I wonder if the underlying improvements in visual language
| learning will allow for even more efficient search. The First
| Fully General Computer Action Model -> https://si.inc/posts/fdm1/
| QubridAI wrote:
| This is a big leap true multimodal search without text
| bottlenecks makes video querying feel finally native and insanely
| practical.
| WatchDog wrote:
| I don't quite understand the 5 second overlap. I assume it's so
| that events that occur over the chunk boundary don't get missed,
| but is there any examples or benchmarking to examine how useful
| this is?
| sohamrj wrote:
| yea, it's so events on a chunk boundary still get captured in
| at least one chunk. i haven't had the chance to do formal
| benchmarks on overlap vs. no-overlap yet. the 5s default is a
| pragmatic choice, long enough to catch most events that would
| otherwise be split, short enough to not add much cost (120
| chunks/hr to ~138). also it's configurable via the --overlap
| flag.
| thegabriele wrote:
| Why just the dash cam?
| sohamrj wrote:
| dashcam is just one of the use cases and the one i tested on.
| but this could theoretically work with any kind of video
| footage like home security footage
| lwarfield wrote:
| Damn, I need to going with my embeddings project. I've currently
| got a prototype for using embeddings (not gemini in my case) for
| making a game that's kinda reverse connections:
|
| collections.lwarfield.dev
___________________________________________________________________
(page generated 2026-03-24 23:00 UTC)