Case Study
On-Device Video Summarization
for Edge ai | Scene Analyzer Agent
Turning Camera Streams into Queryable Scene Intelligence at the Edge
By Doruk Sönmez, M.Sc.
AI Solutions Architect, CTai LABS, a department of Connect Tech Inc. NVIDIA DLI Certified Instructor
Review: Kara Price, Senior Marketing & Events Specialist, Connect Tech Inc, ConnectTech.com
Key Results
- Why Choose CTai LABS for On-Device Video Intelligence: CTai LABS integrates VLMs, LLMs, retrieval, local data storage, and hardware acceleration into an on-device video intelligence pipeline optimized around the customer’s cameras, workloads, and deployment environment
- Scene understanding: Cosmos 3 Edge enables on-device vision reasoning for interpreting activity and events across video streams (NVIDIA, 2026a)
- Hardware-accelerated processing: video decoding, image processing, and ai inference are optimized for local execution on NVIDIA Jetson, helping make long-duration and multi-stream video analysis practical at the Edge
- Privacy: all data, video, summaries, and metadata stay on-premises, supporting air-gapped and bandwidth-limited deployments
- Customizability: fine-tune models for domain-specific tasks and match runtime configuration to the target Jetson profile
- Grounded answers: a CA-RAG layer retrieves relevant clips, timestamps, and metadata to ground natural-language responses in locally stored evidence
What Is On-Device Video Summarization?
Connect Tech’s Scene Analyzer Agent (SAA) is an on-device ai video intelligence application built for privacy-sensitive and connectivity-constrained deployments, where sending footage to the cloud is undesirable or impractical. SAA converts live camera streams and recorded video files into structured, searchable scene information, running entirely on NVIDIA® Jetson™ Edge ai platforms.
The application decodes video frames and passes selected video chunks into a Vision-Language Model (VLM) to generate natural-language scene descriptions. Nemotron-based LLM processing evaluates those descriptions against configured alerts. Summaries, alert records, and clips are stored locally, while combined vector and graph retrieval supports search and natural-language Q&A grounded in the stored evidence.
CTai LABS developed Scene Analyzer Agent and optimized it for deployment across multiple NVIDIA Jetson platforms.
Deployment options include Jetson AGX Orin 32GB, Jetson AGX Orin 64GB, Jetson T4000, and Jetson T5000, selected according to the application’s model size, memory requirements, and number of simultaneous video streams.
How On-Device Video Summarization Works
- Stream or file ingestion. SAA receives live RTSP streams, camera feeds, or prerecorded video files and prepares frames for analysis according to the configured sampling and summarization window.
- Frame decoding and batching. Frames are decoded locally and grouped into temporal chunks so the model can infer both visible entities and event progression across the selected time interval.
- VLM inference. The Vision-Language Model generates a scene-level description that captures objects, people, actions, spatial context, and temporal changes. The system targets state-of-the-art VLM families such as the Cosmos 3 family, selected according to the deployment hardware, memory budget, throughput target, and desired intelligence level.
- Summary normalization and tagging. Model outputs are converted into consistent summaries, tags, timestamps, and event metadata so they can be stored, searched, filtered, and correlated across video sources.
- LLM alert and reasoning layer. Nemotron-family LLMs evaluate whether a scene summary matches configured user alerts, such as a person entering a restricted area or a vehicle arriving at a loading bay. The same LLM layer supports natural-language Q&A, evidence synthesis, and GraphRAG-style reasoning over stored scene records.
- Clip capture and local persistence. When an alert is triggered, SAA saves the relevant video clip, summary, timestamp, alert details, and related metadata in a secure on-device database.
- Hybrid retrieval and grounded response generation. Retrieval-augmented approaches can focus video question answering on the most relevant portions of long-form video before response generation (Gia et al., 2025). Stored summaries are indexed for semantic similarity and as graph relationships among people, objects, locations, and actions. On a question, SAA retrieves relevant evidence through both paths and grounds the response in those local summaries, timestamps, clips, and metadata.
Figure 1: Scene Analyzer Agent seven-step pipeline, from video ingestion and VLM inference through alert evaluation, clip capture, and retrieval.
What This Integration Accomplishes
SAA is positioned as a scalable Edge ai application rather than a single-board-only workload, and it preserves the same functional architecture across every deployment profile. Jetson AGX Orin 32GB in Super Mode is the efficiency-optimized profile, targeting the best intelligence per watt and per memory budget while preserving live stream analysis, alerting, clip capture, local storage, CA-RAG retrieval, and graph-based reasoning. Jetson AGX Orin 64GB adds memory and operational headroom for larger contexts or more feeds.
Jetson T5000 represents the highest-performance configuration in the current SAA deployment range, built on NVIDIA Blackwell architecture with 128GB of unified memory, enabling larger models, higher concurrency, and stronger multi-stream scaling. Jetson T4000 provides a Thor-class option below the T5000 tier, and Jetson T3000 provides a future path for high-performance, power-efficient deployment in smaller-footprint Edge systems when modules become available in Q1 2027.
Underpinning these deployment profiles is a layered optimization pipeline that makes multi-model Edge ai practical on constrained systems. That layered design reflects the CTI EdgeAI Stack approach, aligning Edge compute, accelerated ai software, memory, video processing, model serving, data storage, and deployment requirements as one integrated system. Video decoding and other supported media operations are offloaded to dedicated hardware where applicable, image conversion and scaling use accelerated processing paths, unnecessary visualization stages can be disabled in production, and lightweight serving or quantized model formats can reduce runtime memory overhead (NVIDIA, 2026b).
Every generated summary becomes part of two complementary indexes: a vector index for semantic similarity, and a graph database that captures explicit relationships between people, objects, locations, and actions. This combination allows SAA to answer both meaning-based questions and relationship-based questions while preserving a clear evidence trail back to source clips and timestamps.
Benefits for Customers
Reduced review burden
SAA reduces the amount of manual video review required by converting long-duration footage into time-indexed summaries, alert records, and searchable clips. Operators can query by event, object, location, time window, or semantic description instead of scanning raw footage frame by frame, which matters most for multi-camera environments, archived footage analysis, remote sites, and investigation workflows.
Local data control
the full analysis pipeline, including video ingestion, hardware-accelerated decode, VLM inference, LLM alert evaluation, clip creation, embeddings, graph records, and Q&A outputs, can remain on the deployed edge system. This helps customers maintain control over sensitive video data and operational events without depending on cloud services, and it supports environments with limited bandwidth, intermittent connectivity, strict data-governance requirements, or air-gapped deployment constraints.
Evidence-grounded answers
SAA’s CA-RAG layer combines semantic retrieval with graph relationships to ground responses in locally stored video evidence. SAA searches the on-device database for semantic matches and graph relationships for entity-action-location context, then synthesizes an answer from the relevant summaries, timestamps, clips, and metadata, improving traceability for incident review, compliance support, and forensic search.
Improved operational awareness
By continuously summarizing live streams and evaluating those summaries against configurable alert policies, SAA can surface relevant events without requiring an operator to continuously watch every stream. This supports event-driven monitoring across security, industrial, logistics, smart-city, drone, mining, and agriculture deployments.
Domain-specific and hardware-specific adaptation
SAA can be tuned at both the application and deployment layers. VLM prompts, Nemotron-based LLM behavior, alert rules, summary schemas, graph extraction logic, and Q&A behavior can be customized for customer-specific terminology, objects, locations, and risk conditions, while the serving runtime, model selection, quantization strategy, and compute-offload configuration can be matched to the target Jetson profile.
Applications of Scene Analyzer Agent
- Retail analytics and loss prevention. Summarize shopper movement, shelf interactions, stockroom access, and after-hours behaviour while keeping footage and derived intelligence on premises.
- Warehouse and manufacturing operations. Configure the system to identify and summarize safety-related events, restricted-zone entry, forklift or pallet movement, dock activity, truck arrivals, and workflow-compliance events using configurable alert logic.
- Campus and facility security. Maintain a searchable local event memory for visitor activity, after-hours access, incident review, and clip-based investigation workflows.
- Transportation and logistics. Index activity around loading bays, shipping yards, containers, vehicles, baggage-handling zones, and cargo areas to support investigation, accountability, and operational visibility.
- Smart cities and traffic environments. Analyze pedestrian movement, vehicle interactions, near-miss events, collisions, road activity, and public-space patterns through local Edge processing.
- Drones and autonomous inspection. Process aerial video streams from UAVs to summarize inspection passes, detect objects or anomalies, identify changes across repeated flights, and support search, surveillance, infrastructure inspection, and perimeter monitoring with local Edge inference.
- Mining and heavy industrial sites. Analyze activity around haul roads, loading zones, pits, conveyors, equipment yards, and restricted areas to support safety monitoring, asset tracking, and compliance workflows in remote or connectivity-limited environments.
- Agriculture and precision farming. Adapt the system to summarize field, greenhouse, livestock, and equipment video and surface relevant conditions or changes for review.
Ready to Add On-Device Video Intelligence to Your Deployment?
Bring your video intelligence use case to CTai LABS, Your Physical ai Integration Partner. Our team can integrate and optimize VLM inference, LLM reasoning, alert logic, retrieval, local storage, and hardware acceleration around your cameras, data requirements, and target NVIDIA Jetson platform. Scene Analyzer Agent provides a starting point for turning live and recorded video into locally processed, searchable scene intelligence without requiring cloud connectivity.
ABOUT THE AUTHOR
Doruk Sönmez, M.Sc.
AI Solutions Architect, CTai LABS
Doruk is an AI Solutions Architect at CTai LABS, the Physical AI and Edge AI services division of Connect Tech Inc., an NVIDIA Elite Partner. An NVIDIA DLI Certified Instructor, he specializes in deploying vision-language models, agentic AI workflows, and accelerated video pipelines on NVIDIA Jetson platforms.
Sources and Frequently Asked Questions
Sources
Connect Tech. (2026, July 15). Connect Tech announces support for new NVIDIA Jetson T3000 and T2000 modules.
https://connecttech.com/2026-07-jetson-t3000-announcement/Gia, B. T., Le, K., Do, T., Mai, T.-D., Ngo, T. D., Le, D.-D., & Satoh, S. (2025). VRAG: Retrieval-Augmented Video Question Answering for Long-Form Videos. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 3689–3698.
https://openaccess.thecvf.com/content/CVPR2025W/IViSE/html/Gia_VRAG_Retrieval-Augmented_Video_Question_Answering_for_Long-Form_Videos_CVPRW_2025_paper.htmlNVIDIA. (2026a, July 15). Japan’s robotics and manufacturing leaders build on NVIDIA Cosmos to advance Physical AI frontier. NVIDIA Newsroom.
https://nvidianews.nvidia.com/news/japans-robotics-and-manufacturing-leaders-build-on-nvidia-cosmos-to-advance-physical-ai-frontierNVIDIA. (2026b, August 11). NVIDIA JetPack 7.2.1 adds agentic video skills and T3000 emulation. NVIDIA Technical Blog.
https://developer.nvidia.com/blog/nvidia-jetpack-7-2-1-adds-agentic-video-skills-and-t3000-emulation/Ready to Build Smarter?
Let’s create the intelligent Edge AI solution that moves your business forward.