Cosmos 3 is an omnimodal world model for Physical AI that unifies understanding, generation, simulation, and action across language, images, video, audio, and robot actions in a single architecture.
A flexible camera-controlled generative model for sparse-images-to-video, image-to-video, text-to-360-video, and image-to-sparse-images generation. The project did not have a paper or public release due to company policy.
MLP-Splatting decomposes scenes into a few object-centric light-field primitives, each an independent compact MLP with localized spatial support, enabling photorealistic novel-view synthesis and interactive object-level editing from RGB supervision alone.
KV-Tracker caches key-value pairs from multi-view geometry transformers to enable real-time 6-DoF pose tracking and online scene reconstruction from monocular RGB, achieving up to 15× speedup and ~27 FPS without drift or catastrophic forgetting.
CausNVS is an autoregressive diffusion model for next novel view synthesis with relative pose encoded attention (CaPE) and efficient KV cache inference, towards real-time world modelling, AR streaming and interactive online generation.
EscherNet is a multi-view conditioned diffusion model for view synthesis. EscherNet learns implicit and generative 3D representations coupled with the camera positional encoding (CaPE), allowing continuous relative camera control between an arbitrary number of reference and target views.
We present vMAP, an object-level real-time mapping system, with each object
represented by a separate MLP neural field model, and object models are optimised in parallel via vectorised training.
We use a quadruped robot to complete a pedestrian-following task in
challenging scenarios. The whole system consists of two modules: the perception and planning module,
relying on the onboard sensors.
We present a novel semantic-aided LiDAR SLAM with loop closure based on LOAM,
named SA-LOAM, which leverages semantics in odometry as well as loop closure detection.
We propose a novel semantic graph based approach for large-scale place
recognition in 3D point clouds. A novel semantic graph representation and a fast and effective graph
similarity network is presented.
We propose a framework to achieve point-wise semantic segmentation for 3D LiDAR point clouds.
Competitions & Projects
Zero123-hf: a diffusers implementation of zero123Star Xin Kong code
A Hugggingface Diffusers (merged)
implementation of original Zero-1-to-3. Zero-1-to-3
is a large-scale diffusion models that can control the camera perspective, enabling zero-shot novel
view synthesis and 3D reconstruction from a single image.
Awesome Point Cloud Place RecognitionStar Xin Kong, Lin Li code
A list of papers about point cloud based place recognition,
also known as loop closure detection in SLAM.
Our team built two fully automatic robots, including
machinery, circuit, control and algorithm. I was responsible for visual servo, localization, navigation and decision-making of robots.
Our team built more than 10 complex automatic or semi-automatic robots.
I was responsible for visual servo, which involves computer vision, RGB-D camera calibration, machine learning,
multithreaded programming, ballistic model modeling, etc.
Our team modeled the practical problems (Managing The Zambezi River) proposed by COMAP into mathematical
models. Through background research, reasonable assumptions and optimization analysis, a solution to the problem was obtained.
Our team modeled the practical problems (Mooring System Design) proposed by CSIAM into mathematical
models. Through background research, reasonable assumptions and optimization analysis, a solution to the problem was obtained.
I was a echelon member of the vision group to help the official team members with Ubuntu
environment building, camera calibration, and computer vision algorithm testing. Thanks to my seniors for their careful guidance!
Our team designed an automatic dustin robot that can catch objects. I was in charge
of Kinect development, RGB-D camera calibration, moving object tracking, and trajectory prediction.
Our team designed and implemented an automatic book sterilizer to protect books
by cleaning up the bacteria and dust in books. Patent No. ZL 2015103334672.
Honors
May. 2021, Sun Youxian (Academician of the Chinese Academy of Engineering) Scholarship.