Computer Vision · Autonomous Driving
Building virtual worlds for safer autonomous driving.
I am a Ph.D. candidate in Transportation Engineering at Tongji University, advised by Prof. Ying Ni and co-advised by Prof. Haotian Shi.
My research lies at the intersection of 3D vision, neural simulation, and end-to-end autonomous driving. I build realistic, interactive environments for systematic evaluation, diagnosis, and improvement of autonomous driving systems.
Current frame · 2026.09
Now
Sim-to-real consistency for neural driving simulation.
Evaluation and diagnosis methods for end-to-end driving systems.
Research conversations around 3D vision and autonomous driving safety.
Updates
News
Latest AETHER, our multimodal closed-loop autonomous-driving simulator, is now open source.
Our paper DecoupleGS was accepted to ECCV 2026.
Two papers were published at IEEE ITSC 2025.
Selected work
Publications
Research on interactive simulation, safety-critical scenarios, and evaluation for autonomous driving.
Peer-reviewed
2025–2026
DecoupleGS: Interactive 3D Gaussian Splatting for End-to-End Autonomous Driving Testing
European Conference on Computer Vision (ECCV), 2026
Abstract
End-to-end driving systems need closed-loop testing environments that are simultaneously photorealistic, interactive, and fast enough for online control. DecoupleGS meets these requirements by separating a persistent high-fidelity background from reusable dynamic agents represented in object-centric canonical coordinates, then compositing them through a unified 3D Gaussian rasterizer. Three targeted modules resolve the conflicts introduced by dynamic scene composition: perceptual pruning and vector quantization compress traffic assets; map-guided registration aligns agent trajectories and road contact in metric space; and proxy-based relighting transfers local illumination and contact shadows without online neural inference. Together, they support controllable multi-agent sensor simulation while preserving geometric and photometric consistency.
- System designDecouples static infrastructure from manipulable canonical agents and renders both streams together with physically consistent occlusion.
- EvaluationEvaluated on nuScenes and PandaSet scenes with 3DRealCar assets, plus UniAD and VAD open- and closed-loop testing.
- Key resultRuns at 45 FPS in the simulator comparison, reaches 0.884 Driving Score and 0.956 Route Completion, and scales to 50 agents.
@article{li2026decouplegs,
title={DecoupleGS: Interactive 3D Gaussian Splatting for End-to-End Autonomous Driving Testing},
author={Li, Siying and Ni, Ying and Sun, Jie and Sun, Jian and Shi, Haotian},
journal={arXiv preprint arXiv:2608.01761},
year={2026}
}
VRU-Centric Hazardous Scenario Detection via Monocular Spatiotemporal Feature Fusion
2025 IEEE 28th International Conference on Intelligent Transportation Systems, pp. 4685–4690
Abstract
Hazardous interactions between vehicles and vulnerable road users often emerge through subtle spatial and motion cues that are difficult to identify from monocular video. VRU-HazardNet turns these cues into an explicit risk score with a two-stream pipeline: MonoTTA and AB3DMOT recover tracked monocular 3D geometry, SEA-RAFT and ResNet-50 encode motion, and a causally masked temporal Transformer fuses both streams. Mean-pooled temporal features pass through a calibrated risk head with temperature scaling and Focal Loss, improving confidence stability and sensitivity to rare hazards. The work also introduces VRUHI, a dedicated benchmark for proactive vehicle–VRU risk assessment rather than post-event recognition.
- ArchitectureUses frozen 3D detection, tracking, and optical-flow front ends with a 4-layer, 8-head, 512-dimensional Transformer for temporal risk reasoning.
- DatasetVRUHI contains 6,000 one-second, 25-frame D2-City clips—2,000 hazardous and 4,000 safe—with severity, interaction, and lane-separation annotations.
- EvidenceAchieves 78.37% AUC and 41.03% F1; combining 3D boxes and flow adds 7.09 AUC points over RGB alone and exceeds R(2+1)D by 2.95 AUC points.
@inproceedings{ni2025vru,
title={VRU-Centric Hazardous Scenario Detection via Monocular Spatiotemporal Feature Fusion},
author={Ni, Ying and Li, Siying and Fan, Jialin},
booktitle={2025 IEEE 28th International Conference on Intelligent Transportation Systems (ITSC)},
pages={4685--4690},
year={2025},
doi={10.1109/ITSC60802.2025.11423759}
}
Interactive Adversarial Scenario Generation for Autonomous Driving: A Continual Learning Framework with Safety Constraints
2025 IEEE 28th International Conference on Intelligent Transportation Systems, pp. 4656–4662
Abstract
Safety-critical scenarios are valuable for autonomous-driving validation but rare in natural data, while unconstrained adversarial generation often produces unrealistic or inevitable collisions that leave the tested vehicle no meaningful response. Constrained-Adversarial Policy Optimization (CAPO) introduces adversarial rationality through a two-phase continual-learning framework. Phase I augments independent multi-agent PPO with a learned safety-constraint function that approximates Hamilton–Jacobi reachability, allowing background vehicles to learn realistic task completion and safe interaction. Phase II introduces an autonomous-vehicle expert, freezes the task value function, and optimizes adversarial behavior according to the expert's estimated safety margin. This creates near-conflict interactions that remain challenging and solvable rather than collapsing into unavoidable crashes.
- Constrained policyA dual-value objective combines task advantage with a learned safety value, then uses that constraint to regulate adversarial pressure on the AV expert.
- Training setupTrained in MetaDrive with 20 agents, partial local observations, five random seeds, and 1.5 million steps; compared against Vanilla-IPPO and collision-driven ADV-IPPO.
- ValidationOpen-loop TTC/PET analysis reduces extreme-risk cases, while four closed-loop scenarios preserve corrective space in left-turn, crossing, cut-in, and head-on conflicts.
@inproceedings{fan2025interactive,
title={Interactive Adversarial Scenario Generation for Autonomous Driving: A Continual Learning Framework with Safety Constraints},
author={Fan, Jialin and Ni, Ying and Chen, Yuhang and Li, Siying and Sun, Jie and Sun, Jian},
booktitle={2025 IEEE 28th International Conference on Intelligent Transportation Systems (ITSC)},
pages={4656--4662},
year={2025},
doi={10.1109/ITSC60802.2025.11423840}
}
Research system
AETHER
A continuously evolving engineering platform that turns research ideas into an executable autonomous-driving simulator.
PROJECT // 01
High-fidelity, multimodal closed-loop simulation for end-to-end driving
AETHER is an editable driving-scene pipeline built around 3D Gaussian Splatting. It reconstructs real scenes, inserts and controls traffic actors, synthesizes synchronized RGB and LiDAR observations, and feeds those observations into closed-loop driving and diagnosis. Rather than being a static demo, it serves as the engineering convergence point for the simulation, self-improvement, and evaluation ideas developed across this research.
- 01ReconstructReal-world 3DGS scenes
- 02EditActors, routes, and events
- 03SenseRGB and LiDAR novel views
- 04Roll outEnd-to-end closed loop
- 05RefineDistillation and diagnosis
Editable neural environments
Reconstructs real driving scenes with 3DGS and provides scenario editing for actor placement, trajectory replay, termination logic, and controlled interaction.
Camera–LiDAR synthesis
Combines RGB rendering and enhancement with geometric or learned LiDAR simulation covering range, intensity, ray returns, point clouds, and novel viewpoints.
Policy-in-the-loop testing
Connects synchronized sensor observations to a maintained UniAD adapter, updates the scene from policy actions, and records replayable multimodal rollouts.
Self-improvement and diagnosis
Supports reverse distillation from enhanced observations, semantic-causal diagnosis, and evaluation with HUGSIM-, NAVSIM-, and Bench2Drive-style metrics.
From scene layers to closed-loop driving
The panels follow the simulator from editable scene representation, through multimodal sensing, to policy-in-the-loop evaluation. Playback remains user-controlled.
Decoupled scene composition, multimodal novel-view enhancement, reverse distillation, and beyond-ego diagnosis are brought together in one reproducible pipeline that will continue to evolve.
Where AETHER is—and where it goes next
Open-source release
The simulator code is now public, covering editable 3DGS reconstruction, actor control, RGB and two-mode LiDAR synthesis, Difix restoration, multimodal rollout, reverse distillation, UniAD closed-loop driving, diagnosis, and output evaluation.
Reliable scene-level testing
Current development focuses on reliable testing across reconstructed scenes from public driving datasets. Scenario editing is still centered on one vehicle at a time, while broader traffic composition and scene diversity remain active work.
2D traffic flow × 3D neural rendering
The next stage will connect TESS NG and LimSim for trajectory-level traffic-flow simulation in 2D, then render the evolving traffic state in AETHER's 3D scenes. TESS NG integration is an active collaboration with Jida (济达).
Research agenda
Research
From realistic scene reconstruction to systematic diagnosis of autonomous driving systems.
3D Vision
Reconstructing dynamic, photorealistic driving environments from multimodal observations.
Neural Simulation
Building interactive virtual worlds that respond faithfully to agent behavior.
End-to-End Driving
Evaluating and improving driving systems through closed-loop testing and diagnosis.
Connection map
Papers → AETHER → evaluation
Select a thread to trace how individual studies converge in the simulator and surface as measurable system behavior.
Evaluation
Sim-to-Real Evaluation for Autonomous Driving Simulation
Developing systematic methodologies for measuring how faithfully virtual environments reproduce real-world system behavior, across both open-loop and closed-loop settings.
Diagnosis
Diagnosis of End-to-End Autonomous Driving Systems
Identifying failure-critical temporal windows, traffic participants, and system capabilities to characterize model limitations and safety boundaries.
Understanding
Hazardous Traffic Scenario Understanding
Studying vision-based and 3D spatial reasoning for safety-critical interactions involving vulnerable road users and complex traffic behavior.
Event sensing
Direct Event-Camera Simulation from 3D Gaussians
Developing a simulation pipeline that generates asynchronous event streams directly from Gaussian primitives in reconstructed 3DGS scenes, without using rendered images as an intermediate representation, for closed-loop testing of event-based end-to-end driving systems.
Background
Experience
Education
Ph.D. in Transportation Engineering
College of Transportation, Tongji University
Shanghai, China
B.S. in Mathematics and Applied Mathematics
Guohao School, Tongji University
Shanghai, China
Academic service
Academic service
- ReviewerWiCV at ECCV 2026
- ReviewerECCV Main Conference
- ReviewerIEEE ITSC 2025
- ReviewerTransactions on Computing Science (TCS)
Award
National Third Prize
Huawei Cup China Graduate Mathematical Contest in Modeling, 2024
Patent
Method and System for 3D Gaussian Scene Simulation for End-to-End Autonomous Driving Testing
Ying Ni, Siying Li, Haotian Shi, Jian Sun
Get in touch
Interested in research collaboration?
I am always happy to discuss 3D vision, simulation, and autonomous driving research.