← PROJECTS

VIDEO → MOTION → 3D

Fencing
Zero-Shot 3D

フェンシング映像から人物と剣(剣先を含む)を解析し、3Dに復元する技術デモ。フェンシング専用の教師データを一切作らず、汎用の大規模ビジョン/視覚言語モデル(VLM)だけで剣先を追跡するゼロショット処理。

A technical demo that analyzes fencers and their blades (including the tips) from video and reconstructs them in 3D. Zero-shot: no fencing-specific training data was created; blade tips are tracked using only general-purpose large vision / vision-language models (VLMs).

背景Background

かつて剣先を追跡するには、映像に手動でラベルを付けた大量の教師データを作り、専用モデルを学習させる必要があった。現在は汎用の大規模モデルの性能が向上し、専用データなしでも剣先の位置を推定できるようになりつつある。本プロジェクトはその事実を示す。

Tracking blade tips used to require large hand-labeled datasets and a purpose-built model trained on them. General-purpose large models have now improved to the point where tip positions can be estimated without task-specific data. This project demonstrates that shift.

主張Claim

巨大汎用モデルの役割は、専用モデルの置き換えではない。動画生成AIで合成データを作り、巨大汎用モデルがそれにラベルを付けて教師データを自動生成する。撮影も手動ラベリングも不要になり、そのデータで学習した軽量モデルが超低遅延で毎フレームを追う。巨大モデルは低レートで意味的な判断を担い、軽量モデルが毎フレームを追う。これが当面の現実的な構成だと考えている。

The role of large general-purpose models is not to replace dedicated models. Video generation AI produces synthetic footage; a large general-purpose model labels it, generating training data automatically. Neither filming nor manual labeling is required, and a lightweight model trained on that data tracks every frame at ultra-low latency. The large model makes semantic decisions at low rate; the lightweight model tracks every frame. For now, this is the practical configuration.

3Dデモを開くOpen the 3D demo 自由視点・軌跡・パーティクル Free camera · trails · particles
ORIGINAL VIDEO → 3D元動画とCGを同期再生 / 4つの演出 / 0.5倍速Synchronized source + CG / 4 visual styles / 0.5× speed

見通しOutlook

現状の処理はリアルタイムではない。ただし、リアルタイム化は時間の問題だと考えている。数十ミリ秒以下の超低遅延が求められる用途では専用の軽量な手法が引き続き必要だが、1秒程度の遅延が許容される場面では、巨大な汎用モデルをそのまま使うアプローチが主流になるだろう。

The current pipeline does not run in real time, but real-time operation is only a matter of time. Applications that need ultra-low latency (tens of milliseconds or less) will still call for dedicated lightweight methods, but where a delay of around one second is acceptable, using large general-purpose models directly is likely to become the standard approach.

専用手法についてOn dedicated methods

Rhizomatiksでは2012年から剣先追跡を開発しており、現行システムは手動ラベリングで学習した多段検出ネットワークと、24台の4Kカメラによる多視点3D推定で構成され、東京2020や2026年のWorld Fencing Leagueで実戦使用されている。剣先は4Kでも数ピクセルで、高速に動き、剣がしなる。この条件では高フレームレートの専用検出器+多視点幾何+時系列予測が勝つ。2026年時点なら、RT-DETR/RF-DETR(Keypoint)やYOLO26-poseを少量データで微調整し、剣を剣先・鍔のキーポイント列として扱い、剣先周辺を高解像度で再検出するカスケードと組み合わせる構成が現実的だろう。

Rhizomatiks has developed blade-tip tracking since 2012. The current system uses a multi-stage detection network trained on hand-labeled data and multi-view 3D estimation from 24 4K cameras, and has been used in live competition at Tokyo 2020 and the 2026 World Fencing League. Tips occupy a few pixels even in 4K, move fast, and blades bend; here a high-frame-rate dedicated detector with multi-view geometry and temporal prediction wins. In 2026, a practical build would fine-tune RT-DETR / RF-DETR (Keypoint) or YOLO26-pose on a small dataset, model the blade as a keypoint chain (tip, guard), and add a cascade that re-detects the tip region at high resolution.

この映像向けの追加学習は行っていません。既存の学習済みモデルを利用し、内部では複数段の解析・補正を行います。表示しているのは保存済みの解析結果で、単眼映像からの3D推定です。

No additional training was performed for these videos. The program uses existing pretrained models and several internal stages of analysis and correction. The displayed results are saved analysis outputs: 3D estimates from monocular video.

AI生成したフェンシング映像を使用(実写ではない)。解析結果には、目視確認に基づく手動補正を含みます。このデモではブラウザ上で新しい映像の解析は行いません。軌跡・パーティクル・発光は視覚演出であり、衝突の検出結果ではありません。

The experiment uses AI-generated fencing footage (not real footage). The analysis results include manual corrections based on visual review. This demo does not analyze new footage in the browser. Trails, particles, and glow are visual effects, not detected collisions.