Thanks! I really love your phrasing of 'refusing to hallucinate success' β that's exactly the mindset we aimed for. Glad the philosophy resonates!
Xiangpeng Yang PRO
XiangpengYang
AI & ML interests
diffusion models, video generaiton, video editing
Recent Activity
updated a Space 3 days ago
XiangpengYang/pi0.5 published a Space 3 days ago
XiangpengYang/pi0.5 updated a Space 3 days ago
XiangpengYang/qwengr00tOrganizations
replied to their post 7 months ago
replied to their post 7 months ago
Thanks for the feedback! I've just updated the input video, so the examples should now match the quality shown in the teaser. The current frame rate is set to 8 fps, which is standard for these demos.
Regarding the resolution (fidelity), VideoCoF actually supports arbitrary resolution and arbitrary length. You are welcome to upload your own high-resolution videos to test the performance!
Post
3091
π Introducing VideoCoF: Unified Video Editing with a Temporal Reasoner (Chain-of-Frames)!
Weβre excited to introduce VideoCoF, a unified framework for instruction-based video editing that enables temporal reasoning and ~4Γ video length extrapolation, trained with only 50k video pairs. π₯
π What makes VideoCoF different?
π§ Chain-of-Frames reasoning , mimic human thinking process like Seeing β Reasoning β Editing to apply edits accurately over time without external masks, ensuring physically plausible results.
π Strong length generalization β trained on 33-frame clips, yet supports multi-shot editing and long-video extrapolation (~4Γ).
π― Unified fine-grained editing β Object Removal, Addition, Swap, and Local Style Transfer, with instance-level & part-level, spatial-aware control.
β‘ Fast inference update
π H100: ~20s / video with 4-step inference, making high-quality video editing far more practical for real-world use.
π Links
π Paper: https://arxiv.org/abs/2512.07469
π» Code: https://github.com/knightyxp/VideoCoF
π€ Demo: XiangpengYang/VideoCoF
π§© Models: XiangpengYang/VideoCoF
π Project Page: https://videocof.github.io/
#VideoEditing #DiffusionModels #GenerativeAI #ComputerVision #AI
Weβre excited to introduce VideoCoF, a unified framework for instruction-based video editing that enables temporal reasoning and ~4Γ video length extrapolation, trained with only 50k video pairs. π₯
π What makes VideoCoF different?
π§ Chain-of-Frames reasoning , mimic human thinking process like Seeing β Reasoning β Editing to apply edits accurately over time without external masks, ensuring physically plausible results.
π Strong length generalization β trained on 33-frame clips, yet supports multi-shot editing and long-video extrapolation (~4Γ).
π― Unified fine-grained editing β Object Removal, Addition, Swap, and Local Style Transfer, with instance-level & part-level, spatial-aware control.
β‘ Fast inference update
π H100: ~20s / video with 4-step inference, making high-quality video editing far more practical for real-world use.
π Links
π Paper: https://arxiv.org/abs/2512.07469
π» Code: https://github.com/knightyxp/VideoCoF
π€ Demo: XiangpengYang/VideoCoF
π§© Models: XiangpengYang/VideoCoF
π Project Page: https://videocof.github.io/
#VideoEditing #DiffusionModels #GenerativeAI #ComputerVision #AI
posted an update 7 months ago
Post
3091
π Introducing VideoCoF: Unified Video Editing with a Temporal Reasoner (Chain-of-Frames)!
Weβre excited to introduce VideoCoF, a unified framework for instruction-based video editing that enables temporal reasoning and ~4Γ video length extrapolation, trained with only 50k video pairs. π₯
π What makes VideoCoF different?
π§ Chain-of-Frames reasoning , mimic human thinking process like Seeing β Reasoning β Editing to apply edits accurately over time without external masks, ensuring physically plausible results.
π Strong length generalization β trained on 33-frame clips, yet supports multi-shot editing and long-video extrapolation (~4Γ).
π― Unified fine-grained editing β Object Removal, Addition, Swap, and Local Style Transfer, with instance-level & part-level, spatial-aware control.
β‘ Fast inference update
π H100: ~20s / video with 4-step inference, making high-quality video editing far more practical for real-world use.
π Links
π Paper: https://arxiv.org/abs/2512.07469
π» Code: https://github.com/knightyxp/VideoCoF
π€ Demo: XiangpengYang/VideoCoF
π§© Models: XiangpengYang/VideoCoF
π Project Page: https://videocof.github.io/
#VideoEditing #DiffusionModels #GenerativeAI #ComputerVision #AI
Weβre excited to introduce VideoCoF, a unified framework for instruction-based video editing that enables temporal reasoning and ~4Γ video length extrapolation, trained with only 50k video pairs. π₯
π What makes VideoCoF different?
π§ Chain-of-Frames reasoning , mimic human thinking process like Seeing β Reasoning β Editing to apply edits accurately over time without external masks, ensuring physically plausible results.
π Strong length generalization β trained on 33-frame clips, yet supports multi-shot editing and long-video extrapolation (~4Γ).
π― Unified fine-grained editing β Object Removal, Addition, Swap, and Local Style Transfer, with instance-level & part-level, spatial-aware control.
β‘ Fast inference update
π H100: ~20s / video with 4-step inference, making high-quality video editing far more practical for real-world use.
π Links
π Paper: https://arxiv.org/abs/2512.07469
π» Code: https://github.com/knightyxp/VideoCoF
π€ Demo: XiangpengYang/VideoCoF
π§© Models: XiangpengYang/VideoCoF
π Project Page: https://videocof.github.io/
#VideoEditing #DiffusionModels #GenerativeAI #ComputerVision #AI