We live in an age where artificial intelligence increasingly dominates the internet, making it harder to distinguish real content from fake news. Any picture or video you see today could be AI-generated and the traditional markers that would give it away are fading rapidly. There are tools that help identify if something has been made with AI, and Nvidia — arguably the main beneficiary of the AI race — has just released its own, called the "Synthetic Video Detector" (SVD).
SVD is an Nvidia Inference Microservice (NIM) part of the company's AI for Media Private Access Program, so it's not publicly available to consumers, but a demo exists. Anyhow, SVD's job is simple: detect whether a video is real or if was generated using AI. It can analyze videos at scale, breaking them down frame-by-frame to spot anomalies. It's based on cutting-edge research that won awards at computer vision conference ICCV.
Instead of looking at the video file as a whole, SVD splits it into cropped frames, each carrying a 504x504 resolution. These frames are then passed through two powerful Vision Transformers made by Meta: DINOv2 and DINOv3. A job of a vision transformer is to learn to form patterns without needing human-labeled data. They're commonly used for image classification, image retrieval, object detection, and depth estimation.
As such, once the frames go through these transformers, their distinct spatial features are quickly assessed, and each one is assigned a score between 0 and 1 — 0 representing a fully real image and 1 representing a completely fake image. The scores are tallied at the end to form an average, which tells the user whether the video is real or not based on a percentage score out of 100.
This way, news agencies, broadcasters, and media outlets can authenticate footage they receive much quicker and with better certainty. SVD is even designed to work with the reality of social media compression since videos uploaded online will have their imperfections masked. But the transformers can still detect patterns that the human eye cannot, seeing past surface-level anomalies to instead focus on intrinsic artifacts.
As visible in AI GVD bench above, uncompressed video still delivers the best results with SVD showing an insane 92% accuracy rate. At 15% compression, the model drops down to 87% accuracy, while a 50% compression rate still achieves a very impressive 82% accuracy in detecting AI-generated content. Since this is a microservice, it has exceptional latency as well, processing 1080p video in just 22ms on Nvidia RTX GPUs and 30ms on Nvidia's workstation models.
That being said, SVD requires the NVENC encoder, so datacenter cards like the B100 cannot run it natively. Nvidia said it's already working with Wowza to bring real-time synthetic video detection into livestreaming workflows. A demo version is available to try right now at build.nvidia.com but beware that it takes a long time to process since it happens in the cloud, the max file size limit is only 100MB, and it often just times out.
Follow Tom's Hardware on Google News, or add us as a preferred source, to get our latest news, analysis, & reviews in your feeds.

6 hours ago
7







English (US) ·