IN-VIDEO AI

Video intelligence as one API

Video intelligence that turns your catalog into structured, searchable and AI-ready metadata.

Per-minute pricing, no seat count. $25 in signup credits, no credit card for the free tier.

13AI capabilities
30+languages
sub-secondsearch
per-minutepricing

TRUSTED BY PRODUCT TEAMS SHIPPING VIDEO AT SCALE

Customer logoCustomer logoCustomer logoCustomer logoCustomer logoCustomer logo

Search-ready on upload

Ingest your videos at the same time as encode, and produce multimodal embeddings across visual, audio, speech, and on-screen text in a single vector. Your structured metadata extraction is done when encode is done. No second pipeline. Video API for AI Era.

Multimodal Video AI

AI understands and describes every scene and sequence. Including who appears, what's happening, and understands context, turning video into rich, time-coded text. Transcribe over 50+ languages including diarized conversations, speaker detection and hook worthy soundbites.

Usage-based pricing

Priced beside encoding, per minute. No token billing and no separate AI invoice to forecast.

Faster results

Use our extensive built features to orchestrate outcomes at scale. Or pair the metadata with your RAG pipelines to create custom outputs and workflows. With our great docs, webhooks, and best-in-class support from real engineers, go from idea to launch 10x Faster.

MyClassboard a leader in EdTech uses FastPix to power multimodal video AI at scale.

01 · Extract video metadata

Turn any video into transcripts, chapters, and structured metadata

Video understanding parallel to encodes – not a second job – and when done you get one JSON file: transcript, chapters, people and brands, scenes, and any text on screen. It is all timestamped against the same video. Nothing needs matching later.

And this barely scratches the surface. See our docs for more.

1{
2 "success": true,
3 "data": {
4 "status": "COMPLETED",
5 "result": {
6 "scenes": [
7 {
8 "startTime": 0.0,
9 "endTime": 1.478,
10 "description": "A computer screen displaying a chat application with a list of steps for modifying the application."
11 },
12 {
13 "startTime": 1.478,
14 "endTime": 3.44,
15 "description": "A woman in a blue dress is seated, smiling and speaking. The background includes modern wall decor."
16 },
17 {
18 "startTime": 3.44,
19 "endTime": 8.12,
20 "description": "Close-up of fingers typing lines of Java code on a mechanical keyboard with RGB backlighting."
21 }
22 ]
23 }
24 }
25}
JSON

AI Video Scene Analysis provides shot and scene boundary detection with per-scene visual description. By analyzing visual, audio and narrative elements across the video timeline, it generates metadata that powers downstream workflows like video RAGs, recommendations, ad or quiz placements, nudges and more. No more manual tagging. All this in same pass as encoding.

02 · READY TO USE

AI video moderation, clipping and search on one analysis

FastPix analyzes every video you upload. These four features read that analysis, so there is nothing extra to configure and no second vendor to add. They are also the proof: if our own clipping and search run on this data, your pipeline will run on it too.

Turn long videos into shorts instantly

  • Get a week of social clips out of one long recording without opening an editor.
  • Our AI Video clipping agent understands, segments, scores and renders clips, highlight detection ranks each one on hook, pacing and narrative.
  • Every clip ships with auto-generated captions and a title, ready to post.
Read the guide

03 · BUILD YOUR OWN

Build multimodal video RAG, recommendations, and feed ranking

The same JSON goes to your warehouse. What you do with it is not something we need to have shipped first.

How do you build RAG over video?

Multimodal RAG over video

Your archive already has the answers, it just needs indexing first. Multimodal RAG over video retrieves the right moments, then lets a model answer from them. Transcript, scenes and entities are your chunks, stored wherever you already run semantic retrieval.

  • Answer 'when did we announce the price change' across 400 recorded calls
  • Ground a support bot in the product demos, not just the docs
  • Chunk a two-hour webinar into retrievable moments instead of one blob
Recommend by what is in the video

Recommendations

Titles and tags are a thin signal. What is on screen is a thick one. Scene and entity data lets a recommendation engine match on what happens in the video, so two clips about the same thing look related even when their metadata does not.

  • Surface the other three videos that feature the same product
  • Match on what happens in the scene, not what the uploader typed
  • Relate two videos whose metadata has nothing in common
Rank a feed before you have behavior data

Feed ranking

A new upload starts blind, because feed ranking runs on clicks and dwell time. Scene, object and entity signals give the ranker content from the first impression. That is the cold start problem, handled at ingest.

  • Rank a video on its first impression, before anyone clicks
  • Feed scene and entity signals into the ranker you already run
  • Stop good uploads dying in the cold-start hole
Decide where the ads go, and what plays there

Contextual ads & ad breaks

Most ad decisions never look at what is actually playing. Scene boundaries give you ad break detection, and scene descriptions map to IAB categories for contextual ad targeting. SCTE-35 markers land in the stream at the breaks, ready for the VAST or VMAP response your ad server returns. The decisioning stays yours.

  • Cut breaks at scene boundaries instead of every eight minutes
  • Map a scene to an IAB category before the ad call goes out
  • Write SCTE-35 markers your ad server can act on

BUILD VS BUY

The same video AI, assembled on AWS or Google Cloud

Most teams evaluating this are choosing between one API and a stack they wire together. Here is the honest shape of that choice.

CapabilityOn AWSOn Google CloudOn FastPix
Transcript with speaker labelsTranscribe (+ diarization)Speech-to-Text (+ diarization)One flag
Named entitiesComprehendNatural Language APIOne flag
On-screen text (OCR)Rekognition text detectionVideo Intelligence text detectionOne flag
Scene and object detectionRekognition Video segments + labelsVideo Intelligence shot change + labelsOne flag
Chapters and summaryBedrockGemini on Vertex AIOne flag
Semantic searchOpenSearch + embeddingsVertex AI Vector SearchOne flag
Clipping and reframeMediaConvertTranscoder APIOne flag
ModerationRekognition content moderationVideo Intelligence explicit contentOne flag
OrchestrationStep Functions + LambdaCloud WorkflowsOne flag

Tech specs

What In-Video AI handles.

Features, languages, output formats, integration patterns.

Transcription languages

English
English
Spanish
Spanish
Portuguese
Portuguese
French
French
German
German
Hindi
Hindi
Tamil
Tamil
Telugu
Telugu
Bahasa
Bahasa
Vietnamese
Vietnamese
Thai
Thai
Arabic
Arabic
30+ total
30+ total

Caption formats

WebVTT
WebVTT
SRT
SRT
TTML
TTML
Burn-in subtitles
Burn-in subtitles

Chapter detection

Visual scene boundaries
Visual scene boundaries
Audio segment boundaries
Audio segment boundaries
Title generation
Title generation
Timestamp accuracy ~1 second
Timestamp accuracy ~1 second

Search

Conversational queries
Conversational queries
Returns timestamp + context
Returns timestamp + context
Cross-asset library search
Cross-asset library search
Sub-second response
Sub-second response

Summary

Markdown output
Markdown output
100-300 words configurable
100-300 words configurable
Topic tags included
Topic tags included
Multi-language summaries
Multi-language summaries

Moderation

NSFW classifier
NSFW classifier
Tunable threshold
Tunable threshold
Per-frame flags
Per-frame flags
Configurable policy categories
Configurable policy categories
Webhook on flag
Webhook on flag
Audit log
Audit log

NER

People
People
Organizations
Organizations
Places
Places
Products
Products
Dates
Dates
Custom entity types
Custom entity types

Integration

Inline (asset.create flag)
Inline (asset.create flag)
Webhook delivery
Webhook delivery
Editable in dashboard
Editable in dashboard
API for re-processing
API for re-processing

Questions developers ask

In-Video AI questions, answered.

  • What is a video understanding API?

    It turns a video file into structured data you can query. FastPix returns transcripts with speaker diarization, auto chapters, named entities, scene and object detection, and on-screen text as JSON on a webhook. You set flags on the upload call rather than running a second pipeline after encoding.

  • Can I build RAG over video with this?

    Yes, for the extraction and chunking half of the pipeline. Timestamped diarized transcripts, chapter boundaries, OCR and detections are exactly what a retrieval layer ingests, and chapters give you chunk boundaries that do not cut mid-sentence. You still own the embeddings, the vector store and the model.

  • How is this different from running my own model?

    You do not pick or maintain the model. Set a flag at upload and the output is part of the asset. We tune and re-benchmark the underlying models, and billing stays per minute alongside encoding.

  • What is an advanced summary?

    A structured, meeting-style summary of a recording. One request returns an overview, key decisions, action items with owners, sentiment and speaker talk time as JSON, every item timestamped. The Notes Agent produces the same output automatically.

Pricing

Per-minute AI processing.

Inline with encoding. Same per-minute billing model. See full pricing.

TRANSCRIPTION + CAPTIONS

Per minute transcribed.

$0.048/ minute

30+ languages with VTT and SRT export. Same rate for live auto-generated subtitles.

  • 30+ languages
  • VTT + SRT export
  • Speaker diarization

SEARCH + SUMMARY + NER

Per minute analyzed.

$0.0035/ minute

Markdown summaries, structured entities, video chapters, and a conversational search endpoint. Same per-minute rate across NER, chapters, and summary.

  • Conversational search endpoint
  • Markdown summary
  • Named-entity recognition + chapters

MODERATION + SCENE DETECTION

Per minute moderated.

$0.10/ minute

Tunable NSFW / profanity classifier with audit-log-grade webhooks.

  • NSFW + profanity detection
  • Threshold tuneable per workload
  • Webhook + audit log per decision