IN-VIDEO AI
Video intelligence as one API
Video intelligence that turns your catalog into structured, searchable and AI-ready metadata.
Per-minute pricing, no seat count. $25 in signup credits, no credit card for the free tier.
TRUSTED BY PRODUCT TEAMS SHIPPING VIDEO AT SCALE






Search-ready on upload
Ingest your videos at the same time as encode, and produce multimodal embeddings across visual, audio, speech, and on-screen text in a single vector. Your structured metadata extraction is done when encode is done. No second pipeline. Video API for AI Era.
Multimodal Video AI
AI understands and describes every scene and sequence. Including who appears, what's happening, and understands context, turning video into rich, time-coded text. Transcribe over 50+ languages including diarized conversations, speaker detection and hook worthy soundbites.
Usage-based pricing
Priced beside encoding, per minute. No token billing and no separate AI invoice to forecast.
Faster results
Use our extensive built features to orchestrate outcomes at scale. Or pair the metadata with your RAG pipelines to create custom outputs and workflows. With our great docs, webhooks, and best-in-class support from real engineers, go from idea to launch 10x Faster.
MyClassboard a leader in EdTech uses FastPix to power multimodal video AI at scale.
01 · Extract video metadata
Turn any video into transcripts, chapters, and structured metadata
Video understanding parallel to encodes – not a second job – and when done you get one JSON file: transcript, chapters, people and brands, scenes, and any text on screen. It is all timestamped against the same video. Nothing needs matching later.
And this barely scratches the surface. See our docs for more.
1{2 "success": true,3 "data": {4 "status": "COMPLETED",5 "result": {6 "scenes": [7 {8 "startTime": 0.0,9 "endTime": 1.478,10 "description": "A computer screen displaying a chat application with a list of steps for modifying the application."11 },12 {13 "startTime": 1.478,14 "endTime": 3.44,15 "description": "A woman in a blue dress is seated, smiling and speaking. The background includes modern wall decor."16 },17 {18 "startTime": 3.44,19 "endTime": 8.12,20 "description": "Close-up of fingers typing lines of Java code on a mechanical keyboard with RGB backlighting."21 }22 ]23 }24 }25}AI Video Scene Analysis provides shot and scene boundary detection with per-scene visual description. By analyzing visual, audio and narrative elements across the video timeline, it generates metadata that powers downstream workflows like video RAGs, recommendations, ad or quiz placements, nudges and more. No more manual tagging. All this in same pass as encoding.
1{2 "success": true,3 "data": {4 "status": "COMPLETED",5 "result": [6 {7 "speaker": "speaker_0",8 "startSec": 0.34,9 "endSec": 10.52,10 "text": "Hey everyone. Today I'll show you an overview of our platform here at FastPix.io. To start with, FastPix is an all-in-one API solution for building video-centric products",11 "words": [12 { "word": "Hey", "startSec": 0.34, "endSec": 0.44 },13 { "word": "everyone.", "startSec": 0.52, "endSec": 1.18 },14 { "word": "Today", "startSec": 1.22, "endSec": 1.46 },15 { "word": "I'll", "startSec": 1.5, "endSec": 1.61 },16 { "word": "show", "startSec": 1.66, "endSec": 1.82 },17 { "word": "you", "startSec": 1.88, "endSec": 1.92 },18 { "word": "an", "startSec": 1.96, "endSec": 2.02 },19 { "word": "overview", "startSec": 2.1, "endSec": 2.53 }20 ]21 },22 {23 "speaker": "speaker_1",24 "startSec": 43.04,25 "endSec": 52.98,26 "text": "For our product, we tried building our own video pipeline, and that was a nightmare. Just understanding the basics like global uploads to encoding to CDN took months.",27 "words": [28 { "word": "For", "startSec": 43.04, "endSec": 43.12 },29 { "word": "our", "startSec": 43.16, "endSec": 43.28 },30 { "word": "product,", "startSec": 43.32, "endSec": 43.72 },31 { "word": "we", "startSec": 44.0, "endSec": 44.08 },32 { "word": "tried", "startSec": 44.14, "endSec": 44.4 },33 { "word": "building", "startSec": 44.46, "endSec": 44.76 }34 ]35 }36 ]37 }38}Attributed transcripts API returns automatic speech recognition (ASR) with speaker diarization and word-level timestamps. So every quote is traceable to the second it was said. Transcription covers 50+ languages. Built for meetings, interviews, podcasts and panels.
1{2 "chapters": [3 {4 "chapter": "1",5 "startTime": "0.0",6 "endTime": "13.0",7 "title": "Introduction and Mission",8 "summary": "The speaker introduces FastPix.io as an all-in-one API solution for building video-centric products and discusses their mission to increase the number of successful video products."9 },10 {11 "chapter": "2",12 "startTime": "13.0",13 "endTime": "137.0",14 "title": "Challenges of Building Video Products",15 "summary": "The speaker shares the challenges faced by founders in building video products, including the complexity of video pipelines, delays in feature launches, and lack of visibility into video performance."16 },17 {18 "chapter": "3",19 "startTime": "137.0",20 "endTime": "220.0",21 "title": "Benefits of FastPix",22 "summary": "The speaker highlights the benefits of FastPix, such as 400% faster video workflow development, RESTful APIs and SDKs, transparent pricing, and free credits for experimentation."23 }24 ]25}Auto-generate video chapters titled markers with timestamps written straight to the player or manifest. Add chapters to a video automatically and ship a navigable two-hour video without an editor marking sections by hand, lifting completion rate and cutting drop-off in the first thirty seconds.
1{2 "namedEntities": [3 { "entity": "FastPix.io", "category": "Organization" },4 { "entity": "FastPix", "category": "Product" },5 { "entity": "video pipeline", "category": "Technology" },6 { "entity": "global uploads", "category": "Topic" },7 { "entity": "encoding", "category": "Topic" },8 { "entity": "CDN", "category": "Topic" },9 { "entity": "video transcoding", "category": "Topic" },10 { "entity": "video re-engineering", "category": "Topic" },11 { "entity": "analytics", "category": "Topic" },12 { "entity": "AI", "category": "Topic" },13 { "entity": "on-demand video", "category": "Topic" },14 { "entity": "live streaming", "category": "Topic" },15 { "entity": "real-time quality of experience metrics", "category": "Topic" },16 { "entity": "video player", "category": "Topic" },17 { "entity": "InVideo AI APIs", "category": "Topic" }18 ]19}Named entity recognition across the transcript and on-screen text, timestamped for people, orgs, products and places. Video entity extraction that detects brands in video pull every mention of a brand, person or product across an entire library without watching any of it, counting sponsor appearances and tagging hours of inventory in a single run.
1# Summary23FastPix.io is an all-in-one API solution designed to simplify the development of video-centric products and features. The platform offers on-demand video and live streaming APIs, video data products for real-time quality metrics, a customizable video player with analytics, and InVideo AI APIs for content monetization. It aims to accelerate development by 400% compared to DIY or cloud solutions, offering RESTful APIs, SDKs, and transparent pricing. The platform's success is demonstrated by a customer who reduced their time to market by 96% and cut their total cost of ownership by 94%.45Tags: video-api · live-streaming · video-data · in-video-aiAn abstractive video summary returned as text from the API an AI video summarizer that auto-generates a description no one had to watch the video to write. Summarize a video automatically the moment encoding finishes, cutting time from upload to publish and putting usable metadata across a far larger share of the catalog.
1{2 "success": true,3 "data": {4 "result": {5 "meetingSummary": {6 "title": "FastPix.io Platform Overview",7 "overview": "The meeting introduced FastPix.io as an all-in-one API platform for building video-centric products, covering streaming, analytics, AI, and customizable playback tools."8 },9 "keyDecisions": [],10 "detailedNotes": [11 {12 "topic": "FastPix platform overview and mission",13 "startTime": "00:00:00",14 "summary": "Introduced FastPix.io as an all-in-one API platform for building video-centric products and adding video features to existing products.",15 "points": [16 "Positioned as an all-in-one API solution for video products and features.",17 "Mission: increase the number of successful video products in the world."18 ]19 },20 {21 "topic": "Challenges teams face building video infrastructure",22 "startTime": "00:00:20",23 "summary": "Highlighted how difficult it is to build video infrastructure from scratch, even for experienced teams.",24 "points": [25 "Many founders underestimate the complexity until they start implementation.",26 "Global uploads, encoding, and CDN setup can take months.",27 "Feature launches delayed by repeated video re-engineering."28 ]29 },30 {31 "topic": "FastPix product capabilities",32 "startTime": "00:01:26",33 "summary": "Presented as a full-stack video platform combining streaming, analytics, and AI capabilities.",34 "points": [35 "On-demand video and live streaming APIs.",36 "Video data product captures real-time quality-of-experience metrics.",37 "In-video AI APIs enable search, organization, and analysis of videos."38 ]39 }40 ],41 "actionItems": [],42 "openQuestions": [],43 "keyMetrics": [44 { "label": "Founders surveyed", "value": "20+" },45 { "label": "Build speed gain", "value": "400% faster" },46 { "label": "Free credits", "value": "$25" },47 { "label": "Time reduction", "value": "96% less time" },48 { "label": "Cost reduction", "value": "94%" }49 ],50 "datesAndTimes": [],51 "sentiment": {52 "positive": 58,53 "neutral": 34,54 "negative": 855 },56 "speakerTalkTime": [57 { "speaker": "speaker_0", "percentage": 83, "wpm": 161 },58 { "speaker": "speaker_1", "percentage": 9, "wpm": 183 },59 { "speaker": "speaker_2", "percentage": 8, "wpm": 182 }60 ]61 },62 "status": "COMPLETED"63 }64}A video understanding API that takes your own prompt or schema and returns structured sections, topics and key points as JSON. Ask your own question of a video and get structured video metadata back instead of one canned paragraph extract structured data from video in a single call that replaces a custom pipeline and the engineering days per integration.
02 · READY TO USE
AI video moderation, clipping and search on one analysis
FastPix analyzes every video you upload. These four features read that analysis, so there is nothing extra to configure and no second vendor to add. They are also the proof: if our own clipping and search run on this data, your pipeline will run on it too.
Turn long videos into shorts instantly
- Get a week of social clips out of one long recording without opening an editor.
- Our AI Video clipping agent understands, segments, scores and renders clips, highlight detection ranks each one on hook, pacing and narrative.
- Every clip ships with auto-generated captions and a title, ready to post.
Reframe any video to every aspect ratio
- Turn one 16:9 master into 9:16, 1:1, 4:5 or 16:9 outputs, no manual re-crop per platform.
- AI reads the video's genre, podcast, interview, gameplay, sports, vlog, and applies a layout tuned to it.
- Subject-aware tracking keeps speakers and the action centered as they move across the shot.
- Dynamic layouts frame every version to look native on Shorts, Reels and TikTok.
- Runs in the same encode pipeline, priced per minute.

Find moments in secs, not hours
- Find a moment you can describe but can't name, across a petabyte archive? With no tags applied in advance?
- Our Multimodal AI extracts and combines visuals, audio, languages and in-screen text, into a single embedding. It further identifies shots and scenes with per-scene visual disruptions as JSONs.
- With this underlying video metadata you can now search inside your videos, and find the right moments across videos using natural language.
- All this in the same encode pipeline. Priced per minute.
Moderate user-generated video at scale
- Accept user uploads at volume without a human queue standing between upload and publish.
- AI Video Moderation frame and audio classifiers for nudity, violence, prohibited content, with confidence scores.
- We do NSFW and policy classification per frame and flag frames or audio for your review.
- With tunable threshold from 0.0 to 1.0 and webhook on flag with an audit log, you can finetune the moderation process.
- Slash review cost per hour ingested and time to publish for UGC.

03 · BUILD YOUR OWN
Build multimodal video RAG, recommendations, and feed ranking
The same JSON goes to your warehouse. What you do with it is not something we need to have shipped first.
Multimodal RAG over video
Your archive already has the answers, it just needs indexing first. Multimodal RAG over video retrieves the right moments, then lets a model answer from them. Transcript, scenes and entities are your chunks, stored wherever you already run semantic retrieval.
- Answer 'when did we announce the price change' across 400 recorded calls
- Ground a support bot in the product demos, not just the docs
- Chunk a two-hour webinar into retrievable moments instead of one blob
Recommendations
Titles and tags are a thin signal. What is on screen is a thick one. Scene and entity data lets a recommendation engine match on what happens in the video, so two clips about the same thing look related even when their metadata does not.
- Surface the other three videos that feature the same product
- Match on what happens in the scene, not what the uploader typed
- Relate two videos whose metadata has nothing in common
Feed ranking
A new upload starts blind, because feed ranking runs on clicks and dwell time. Scene, object and entity signals give the ranker content from the first impression. That is the cold start problem, handled at ingest.
- Rank a video on its first impression, before anyone clicks
- Feed scene and entity signals into the ranker you already run
- Stop good uploads dying in the cold-start hole
Contextual ads & ad breaks
Most ad decisions never look at what is actually playing. Scene boundaries give you ad break detection, and scene descriptions map to IAB categories for contextual ad targeting. SCTE-35 markers land in the stream at the breaks, ready for the VAST or VMAP response your ad server returns. The decisioning stays yours.
- Cut breaks at scene boundaries instead of every eight minutes
- Map a scene to an IAB category before the ad call goes out
- Write SCTE-35 markers your ad server can act on
BUILD VS BUY
The same video AI, assembled on AWS or Google Cloud
Most teams evaluating this are choosing between one API and a stack they wire together. Here is the honest shape of that choice.
| Capability | On AWS | On Google Cloud | On FastPix |
|---|---|---|---|
| Transcript with speaker labels | Transcribe (+ diarization) | Speech-to-Text (+ diarization) | One flag |
| Named entities | Comprehend | Natural Language API | One flag |
| On-screen text (OCR) | Rekognition text detection | Video Intelligence text detection | One flag |
| Scene and object detection | Rekognition Video segments + labels | Video Intelligence shot change + labels | One flag |
| Chapters and summary | Bedrock | Gemini on Vertex AI | One flag |
| Semantic search | OpenSearch + embeddings | Vertex AI Vector Search | One flag |
| Clipping and reframe | MediaConvert | Transcoder API | One flag |
| Moderation | Rekognition content moderation | Video Intelligence explicit content | One flag |
| Orchestration | Step Functions + Lambda | Cloud Workflows | One flag |
Tech specs
What In-Video AI handles.
Features, languages, output formats, integration patterns.
Transcription languages
Caption formats


Chapter detection
Search
Summary
Moderation
NER
Integration
Questions developers ask
In-Video AI questions, answered.
What is a video understanding API?
It turns a video file into structured data you can query. FastPix returns transcripts with speaker diarization, auto chapters, named entities, scene and object detection, and on-screen text as JSON on a webhook. You set flags on the upload call rather than running a second pipeline after encoding.
Can I build RAG over video with this?
Yes, for the extraction and chunking half of the pipeline. Timestamped diarized transcripts, chapter boundaries, OCR and detections are exactly what a retrieval layer ingests, and chapters give you chunk boundaries that do not cut mid-sentence. You still own the embeddings, the vector store and the model.
How is this different from running my own model?
You do not pick or maintain the model. Set a flag at upload and the output is part of the asset. We tune and re-benchmark the underlying models, and billing stays per minute alongside encoding.
What is an advanced summary?
A structured, meeting-style summary of a recording. One request returns an overview, key decisions, action items with owners, sentiment and speaker talk time as JSON, every item timestamped. The Notes Agent produces the same output automatically.
Pricing
Per-minute AI processing.
Inline with encoding. Same per-minute billing model. See full pricing.
TRANSCRIPTION + CAPTIONS
Per minute transcribed.
$0.048/ minute
30+ languages with VTT and SRT export. Same rate for live auto-generated subtitles.
- 30+ languages
- VTT + SRT export
- Speaker diarization
SEARCH + SUMMARY + NER
Per minute analyzed.
$0.0035/ minute
Markdown summaries, structured entities, video chapters, and a conversational search endpoint. Same per-minute rate across NER, chapters, and summary.
- Conversational search endpoint
- Markdown summary
- Named-entity recognition + chapters
MODERATION + SCENE DETECTION
Per minute moderated.
$0.10/ minute
Tunable NSFW / profanity classifier with audit-log-grade webhooks.
- NSFW + profanity detection
- Threshold tuneable per workload
- Webhook + audit log per decision
Three ways to get unstuck
Whatever kind of help you need, there is a path.
Engineering support
Talk to a video engineer.
Stuck on an API call, a webhook signature, or a player integration? Reach the engineering team directly. Response within hours, not days.
Contact engineeringIntegration help
Docs, code samples, video tutorials.
Self-serve resources for the most common integrations. Quickstart guides, SDK examples, and detailed playback logs in your dashboard.
Browse the docsSolution architect
Plan the rollout with a human.
New integration, migration off another platform, or a complex multi-tenant build. Book a session with a FastPix solution architect.
Join the Slack communityDeveloper resources
Everything you need to start building.
Five-minute quick-start
Sign up, hit the endpoint, ship.
Quick-start guideFull API reference
Every endpoint, every parameter, every response.
API referenceWebhook reference
Every event FastPix emits, with sample payloads.
WebhooksCode samples
Sample apps and SDK examples on GitHub.
GitHubSlack community
Talk to FastPix engineers and other developers.
Join SlackService status
Real-time uptime and incident reports.
Status page


