FELA · Temporal Grounding Head 19.75M params
Video Needle In Haystack Search
Setup
Where did a given moment happen in a video clip?
Pros
Who else can you call to search hours of video cheaply!?
Usage
These are precomputed results - but real! Give it a read.
A small FELA temporal grounding head sits on top of our frozen FEVER model video
features. Given a natural language query and a video's per frame
features, it predicts the single time span - a start and an end
second - where that moment happens - and eerily well on longer videos.
How can it do VIDEO?
The video is encoded once, offline, into per frame features
by a FEVER encoder (our FELA idea applied to vision tasks). The query text is encoded by an off the shelf embedding system text tower.
The 19.75M parameter head reads both and regresses a single [start, end] span in the video's own
timeline.
What does this demo show?
Each clip below is a real held out test video that the model hadn't seen, and the band you
see is the model's real predicted span for that video's real posed haystack retrieval
question. Nothing here is re run live. The page plays the local
clips/*.mp4 file and overlays the prediction, which was already
computed offline, synced to the video's own currentTime.
If you want to see this make it out of research preview stage - so do we! As you might have guessed
our models have extremely high context limits, hours of video search are possible where other models struggle!
Why can't I play with it?
We wish you could - but we are an ethical AI company and unlike some other so called
public benefit corporations in our space - we actually mean it. We pulled a video dataset
for research purposes only, with noncommercial terms. This is a research preview simply to demonstrate
what is possible. Furthermore, we do a lot of vetting on our datasets to prove lineage and provenance. Paying
artists is of paramount importance for us. Without our model, video data is already hard to search to prove safety, with external confirmation - so we need to be careful!