让视频模型学得更好、更快
Getting video models to learn better, faster

原始链接: https://www.linum.ai/field-notes/data-filtering-gen-video

本摘要概述了 **Linum v3** 数据流水线的演进过程,重点介绍了从脆弱的手动启发式方法向高性能、模型驱动的筛选机制的转变。 作者强调,模型的质量从根本上取决于其训练数据;高质量数据能让模型高效学习,而低质量数据只会浪费算力。其方法经历了三个阶段: 1. **2024 年(启发式方法):** 依赖基于 CPU 的计算机视觉(如 PySceneDetect、Haar 级联分类器)和 H.264 运动矢量。这种方法成本低廉,但过于脆弱、难以解释,且缺乏处理复杂视频概念所需的精度。 2. **2025 年初(微调大模型):** 摒弃人工规则,转而使用经过监督微调(SFT)的视觉语言模型(如 Qwen-2-VL)。通过利用数千个 GPU,他们高效处理了数十亿帧数据,取得了显著更好的效果。 3. **2025 年末(RLVR):** 采用带有可验证奖励的强化学习(RLVR)来训练美学评分器。与标准 SFT 相比,这种“手术式”的方法能够更精细地控制数据质量。 最终,Linum 流水线将数十亿个样本精炼为高质量数据集。其核心经验是:**数据筛选是提升模型性能最关键的杠杆。** 应避免使用“廉价”的启发式方法,并投入资源进行高质量、迭代式的模型驱动筛选,以确保模型有效学习。

Hacker News 最新 | 过往 | 评论 | 提问 | 展示 | 招聘 | 提交 登录 让视频模型学习得更好、更快 (linum.ai) 5 分,由 schopra909 发布于 40 分钟前 | 隐藏 | 过往 | 收藏 | 讨论 | 帮助 指南 | 常见问题 | 列表 | API | 安全 | 法律 | 申请 YC | 联系 搜索:
相关文章

原文

Image and video models have gotten a lot better over the last few years, even though the internals of these models haven't changed much since Stable Diffusion 3.

  1. Data Filtering & Rebalancing: Remove noisy data and resample your data strategically so your model learns more effectively
  2. Data Annotation: Gather better annotations like richer captions, bounding boxes, and font details so that it's easier for your model to disambiguate visual concepts
  3. Synthetic Data Generation: Finetune an ensemble of existing generative models to create training data for which there is little-to-no naturally occurring data (e.g. image editing / reference-conditioning for Nano-Banana style models)

A couple of years ago, the prevailing wisdom across all generative models (be it text, image, audio) was to aggregate as much data as humanly possible for pre-training. Luckily, the field has gotten a lot smarter about this. If you throw a bunch of low-quality data (e.g. heavily compressed JPEGs) into pre-training, your model is going to waste a significant amount of its capacity learning how to mimic this slice of data. If you filter your dataset well, your model will have a lot easier time learning what you want it to learn.

We know this sounds obvious, but it's a lot harder to do in practice.

Today we're going to walk you through how our approach to data filtering has evolved since 2024. And, hopefully we'll save you from a couple of headaches if you end up training your own generative models down the line.

  1. 2024Old-school CV on CPUs

    CPU

  2. Early 2025Finetuned LLMs on GPUs

    GPU

  3. Late 2025Reinforcement Learning

    GPU

  • kept
  • thrown out
  • thrown out by mistake
  • kept but should be thrown out
  • RL rubric

[2024] Filtering on a budget — Traditional CV on CPUs

On the first go around, we decided to push our raw dataset through old-school computer vision algorithms. This way we could get away with a cluster of cheap CPU instances instead of an unholy number of GPUs running a multimodal LLM.

Scene detection

We need to filter down tens of billions of images and videos to create our pre-training dataset. Images don't really require any specific pre-processing, but raw videos do.

Next time you watch a television show or movie, track how often the camera cuts. If you're watching something made in the last twenty years, more likely than not you'll see a cut every 5 seconds. When to cut and how to cut is an authorial decision, not something a generative video model should do arbitrarily. So, we need to slice n' dice our videos on shot boundaries into video clips before we can filter them down.

With our cheapskate CPU-only agenda, we picked up PySceneDetect. At a high level it maintains a rolling window of K-frames and if the K+1 frame has significantly different image statistics, it categorizes the frame as a cut. There's no underlying machine learning model. It runs really fast but struggles with common transitions like dissolves, fades, and jitter cuts (which low key is a huge issue).

Getting to know your data

Whenever you get new data, you should spend a few days reviewing random samples, listing what you'd like to keep and what you'd like to throw out. Ideally, you take the time to draft an ontology of categories within "good" and "bad" and track the relative sizes of these categories.

At some point during the data filtering process, your engineer brain will take over, and you'll spend way too much time tuning the knobs of your heuristics (or LLMs), chasing that "perfect" decision boundary. These notes are going to save you from yourself down the line. They'll give you the facts you'll need to talk yourself out of trying "one more idea", when the answer is clearly "no".

Plus, understanding the shape of the data distribution will really help with dataset rebalancing. Certain categories are overrepresented in the natural distribution of all videos. We need to subsample and suppress this signal, otherwise it will dominate training and our model will struggle to learn the long-tail of people/places/things/actions that we need in order to generate anything.

Sieving out the un-captionable

Generative video models are primarily limited by what we can describe correctly and consistently in words.

For now, we need to filter out clips where the primary "thing" that makes the video clip interesting is un-captionable. Without a crystal clear text description, it's just noise to our text-to-video model.

Text-heavy

For example, we want to filter out text-heavy videos. It's still hard for LLMs to caption motion graphics that are constantly changing on screen.

To do this, we sampled frames from each video and ran a tiny EAST Detector to extract bounding boxes for text. From there, we filtered out text heavy videos based on the percentage of the frames that had text and the percentage of each frame covered in text.

Using a CNN for this task was a good idea, but the specific choice was wrong. In order to run tens of billions of frames on CPUs, we had to resize the frames aggressively. So, a lot of text-heavy samples with small fonts fell through the cracks.

Small text goes undetected, after EAST image pre-processing

Original

Original

EAST input

EAST input

EAST is a pretty old model from 2017. It's small and far from the state of the art on text detection. Getting it to run efficiently on CPUs without cache-thrash and thread oversubscription was a challenge. Even after performance optimizations, it was still the largest bottleneck for this version of the data pipeline.

Indescribable actions

When there's not much happening on the screen (e.g. close-up on a person's face), it's hard to describe the specific action taking place. If there's too much happening (e.g. extremely shaky camera, a soccer match with a bunch of folks moving across the pitch at once), LLMs struggle to caption the clip correctly. We lumped these categories of videos together as "indescribable action" clips to be thrown out.

Videos are typically serialized on disk in a compressed format. Codecs like H.264 reduce file size by storing keyframes and motion vectors that describe how the keyframes change over time, rather than RGB values for each pixel over time.

We used the motion vectors stored within the mp4 files themselves to isolate and filter out the "indescribable action" videos.

  • average_frame_energy: L2-norm of all motion vectors averaged across the video
  • min(sub_clip_average_frame_energy): Split each clip into a variable number of chunks depending on the video's length, calculate average frame energy for each chunk, and take the minimum across these L2-norms

Then came the decision tree:

  • average_frame_energy < 0.1: Throw the clip away. These were essentially static videos (e.g. slideshows, still frames, freeze-frames).
  • average_frame_energy > 25: Throw the clip away. The footage was incredibly chaotic.
  • min(sub_clip_average_frame_energy) < 0.03: Throw the clip away. A portion of the clip has nothing happening (e.g. a fade, transition to a still image in a documentary).
  • Keep everything else.

This works well as a cheap first filter, but it has mediocre recall (i.e., a lot of indescribable action clips are kept in the dataset).

Subsampling talking head video clips

From our initial review of the raw data distribution, it was pretty obvious that talking head clips where folks talk straight to camera were dramatically over-represented. If we let the dataset be, it would have been significantly biased towards this sort of clip. Our model would get disproportionately good at creating them (likely at the expense of others), so we needed to find them and subsample them.

Living in an old-school CV world, we naturally burrowed deeper down the engineering tunnel and introduced additional heuristics. We sampled frames from each video clip and ran a Haar-cascade face detector to extract bounding boxes for faces and calculated two numbers:

  • average_face_frame_energy: L2-norm of motion vectors within face bounding boxes
  • average_background_frame_energy: L2-norm of motion vectors, just in the corners of the frame (as a proxy for background motion)

And from there, another decision tree:

  • Moderate average_frame_energy + average_face_frame_energy >= average_background_frame_energy: Keep the clip. Usually a really good close-up.
  • Low/Moderate average_frame_energy: Subsample these.

[Early 2025] rm -rf — Replacing hand-crafted heuristics with finetuned LLMs

We're starting to sketch a rather complicated decision tree. It's full of lossy proxies that only kind of work, and it's very incomplete.

This approach simply doesn't scale. Every time you have a new idea for a filter you have to re-examine how the new node in the decision tree impacts all the other branches. Everything is intertwined and eventually you end up with a pipeline that's both un-interpretable and uneditable.

When we started in 2024, we were staring down the barrel of tens of billions of samples. Given our limited budget, our gut was to construct the cheapest filters possible. This was fundamentally wrong. Our video model struggled to learn basic motions like guitar strumming after training for several weeks on our 2024 dataset but was able to learn these very actions in less than 24 hours of training, after applying our 2025 filters. Instead of looking for the cheapest filters possible, you should optimize for the best possible filters you can afford.

Migrating to 1000s of GPUs

This brings us to our second takeaway: throw away your "principled" computer vision techniques and adopt black box neural networks wherever you can. There are patterns that humans simply can't describe well, no matter how hard we try. Old school CV methods were the best hand-crafted approximations of their era. They're truly impressive feats of engineering, but a well-trained neural network will learn a non-linear function that will win on precision and recall in 99% of cases.

Once you re-orient yourself around this reality, your job should shift from crafting cheap heuristics to optimizing models for GPU throughput and engineering resilient, parallelizable workloads to run on SPOT instances across providers.

Concretely, we replaced PySceneDetect's heuristics with AutoShot and TransNetV2, accelerating inference with custom CUDA kernels.

Iterative dataset labeling (aka self-consistency is harder than you think)

Finetuning is pretty straightforward thanks to the folks at Unsloth.

The hard part of training LLMs for data filtration is that the categories are always somewhat fuzzy. You'll have to answer questions like:

  • If a sample fits several categories to different degrees, which label do I assign it?
  • Should I simplify my categories, so I can label the data more quickly and consistently? Or, do I need to split my category into pieces to make it clearer?
  • How easy is this concept for the LLM to learn? How much data do I need for each category?

More often than not, the biggest problem you'll run into is one of self-consistency. Over the course of labeling a couple hundred samples, it's only natural that you'll relax your criteria, mislabel samples, and muddy the signal in your dataset.

  1. 1Define your categories

    Write clear definitions. Be as specific as humanly possible. You should already have a draft ontology from your dataset study.

  2. 2Label ~2K samples

    Draw random samples from your dataset and label them. Revise categories as you see fit.

  3. 3Split train and validation

    Hold out a validation set. You'll use it once at the very end to make sure your model generalizes. Don't use it to steer the iterative labeling.

repeat k times

  1. 4Finetune the LLM

    LoRA or fully-finetune. We used Qwen-2-VL-2B for our initial filters; smaller models are sufficiently intelligent for these tasks.

  2. 5Evaluate on the training set

    Find the categories the model struggled with the most. Either you need more data for the category or more likely than not, the category is poorly constructed and needs to be redefined. For aesthetic scoring, mislabeled samples are usually indicative of inconsistent grading on your part.

  3. 6Refine categories and re-label

    Add, drop, or merge categories. Then label again. Stop when you can no longer induce a better decision boundary within your LLM.

In a CMU study from 2025, researchers hired cinematography experts to annotate camera motion in online video clips and train other laypeople to make similar annotations. Even with the criteria in hand, the experts disagreed with "ground truth" ~24% of the time. Only through repeated trials were they able to converge on 96% agreement. It's painful to spend days labeling and re-labeling a dataset, but them's the breaks. At least, we're lucky to live in an era where you only need 2-3K samples to train a good filter.

If you find yourself working on data filtration, we'd recommend you hack together a simple labeling tool like the one we show below. We just slapped together a super simple React app with Supabase to store labels and R2 to store the samples.

Labeling Pro-tips

  • Use Hotkeys: You'll want to label as fast as possible (or you'll go crazy). Make sure you can label via hotkeys and that you can edit hotkey mappings easily within the app itself.
  • Make Datasets Forkable: You'll be taking several turns on your dataset, so it's helpful to have a fork feature, where you seed a new dataset from your old labels. Even better if you can quickly drill down to the training samples your model misclassified. These are especially problematic. You'll need to review them to iterate on your criteria effectively. Plus, you'll want to relabel them first.
  • Add tools for label mapping: As you iterate on your ontology, categories will come and go. So, you'll need to make it easy to assign samples that were labeled A to another category B, as you add, merge, and delete groupings.

Turning LLMs into categorical classifiers

Ultimately, we supervise-finetuned (SFT) Qwen-2-VL-2B to tag:

  • Image Categories: Ugly Product Image, Diagram / Screenshot, Collage, Watermarked, Bad Lighting, Pixelated, Drawing / Illustration, Keep
  • Video Categories: Animation, PoV, Bars, Motion Graphics, Ken Burns, Shaky Camera, Little to No Motion, Weird Transition, Keep

We only kept images that our filter predicted as Drawing / Illustration or Keep. For videos, we retained Animation and Keep clips wholesale, while subsampling PoV.

[Late 2025] Reinforcement Learning with Verifiable Rewards (RLVR) for aesthetic filtering

At the start of the data labeling process, we tried to get extremely specific about the properties of the images and videos that divvied up samples into ugly vs. pretty (e.g. overexposed lighting, muted color grades). We thought it would be easier for the LLM to learn the precise reasons why we considered an image ugly than learn an arbitrary "ugliness score".

Once again, our initial intuition turned out to be wrong. The properties that make a particular sample ugly tend to be correlated; you end up assigning K different aesthetic tags to the same sample. And in turn, this poses two significant challenges:

  • Sparse Data Signal: The combinatorial explosion of tags makes it harder for the model to disentangle the categories, especially with a small dataset of a few thousand labels.
  • Slow, Inconsistent Labeling: It takes a lot longer to label samples (and it's a lot harder to be self-consistent) when you have the cognitive load of weighing several possible tags per sample.
We ended up grading the samples on a scale from 1 to 4 and keeping the samples that our models labeled 3 or 4.
联系我们 contact @ memedata.com