Researchers introduced TRACE, a training-free agent for long-horizon video understanding, together with VES-Bench, a 600-question benchmark that evaluates whether decoded video frames fully cover evidence required for temporal ordering and event counting tasks. VES-Bench defines jointly necessary evidence intervals for 348 public long videos, enabling stricter auditing of which frames models actually inspect, addressing concerns that previous evaluations focused mostly on final answers rather than underlying visual support. Under a shared backbone, TRACE achieves 50.7% audited correctness with efficient frame usage, outperforms uniform decoding at comparable costs, attains 63.5% peak answer accuracy, and remains competitive on Video-MME, LVBench, and LongVideoBench benchmarks.
This update represents a notable development in the Ai sector. Organizations and founders tracking this space should evaluate potential strategic and technical implications on their operations.