Articles
Vol. 1 No. 4 (2026)
Kinematically Constrained Diffusion Policies for Contact-Rich Bimanual Robotic Manipulation
School of Computer Science and Applied Mathematics, University of the Witwatersrand, Johannesburg 2001, South Africa
-
Submitted
-
August 7, 2026
-
Published
-
August 7, 2026
Abstract
The orchestration of bimanual robotic manipulation in contact-rich environments represents one of the most formidable challenges in modern robotics. Traditional control paradigms frequently struggle to accommodate the complex hybrid dynamics inherent in multi-contact scenarios, whereas contemporary data-driven approaches, particularly those leveraging deep reinforcement learning or imitation learning, often generate actions that violate physical kinematic constraints. This paper introduces Kinematically Constrained Diffusion Policies, a novel theoretical and algorithmic framework designed to address these fundamental limitations. By integrating differentiable forward kinematics and joint-limit projection mechanisms directly into the iterative denoising process of a diffusion model, the proposed architecture guarantees the generation of kinematically feasible action trajectories while preserving the expressivity and multimodal action distribution modeling capabilities of diffusion policies. The methodology effectively bridges the gap between high-level generative action planning and low-level compliance control. Extensive evaluation across a suite of complex bimanual tasks, including high-precision assembly, flexible material manipulation, and heavy object transport, demonstrates that the inclusion of structural kinematic priors significantly reduces joint limit violations and erratic end-effector movements. Furthermore, the kinematically constrained framework improves sample efficiency and contact stability compared to unconstrained generative baselines. This study provides a comprehensive theoretical analysis of constrained score-based models in continuous control domains and offers extensive empirical evidence supporting the integration of deterministic physical constraints within stochastic generative policies.
References
- 1. Fox, D., Burgard, W., & Thrun, S. (1997). The dynamic window approach to collision avoidance. IEEE Robotics & Automation Magazine, 4(1), 23–33.
- 2. Liu, Fuqiang, Weiping Ding, Luis Miranda-Moreno, and Lijun Sun. "Error adjustment based on spatiotemporal correlation fusion for traffic forecasting." Information Fusion (2025): 103635.
- 3. Zhang, M., Fang, Z., Wang, T., Lu, S., Wang, X., & Shi, T. (2025). Ccma: A framework for cascading cooperative multi-agent in autonomous driving merging using large language models. Expert Systems with Applications, 282, 127717.
- 4. Liu, C., Zhang, J., Zhang, T., Wang, Y., Zhou, H., & Jin, Q. (2026). HRIBench: Benchmarking Interaction-Centric Human-Robot Collaboration.arXiv preprint arXiv:2607.13056.
- 5. Lv, Qi, et al. "Spatial-temporal graph diffusion policy with kinematic modeling for bimanual robotic manipulation." Proceedings of the Computer Vision and Pattern Recognition Conference. 2025.
- 6. Yang, Y., Ma, X., Li, C., Zheng, Z., Zhang, Q., Huang, G., ... & Zhao, Q. (2021). Believe what you see: Implicit constraint approach for offline multi-agent reinforcement learning. Advances in Neural Information Processing Systems, 34, 10299-10312.
- 7. Peng, Qucheng, Chen Bai, Guoxiang Zhang, Bo Xu, Xiaotong Liu, Xiaoyin Zheng, Chen Chen, and Cheng Lu. "NavigScene: Bridging Local Perception and Global Navigation for Beyond-Visual-Range Autonomous Driving." arXiv preprint arXiv:2507.05227 (2025).
- 8. Yang, Y., Hu, H., Li, W., Li, S., Yang, J., Zhao, Q., & Zhang, C. (2023, June). Flow to control: Offline reinforcement learning with lossless primitive discovery. In Proceedings of the AAAI Conference on Artificial Intelligence (Vol. 37, No. 9, pp. 10843-10851).
- 9. Brohan, A., Chebotar, Y., Finn, C., Hausman, K., Herzog, A., Ho, D., et al. (2023). RT-2: Vision-language-action models transfer web knowledge to robotic control. In Proceedings of the Conference on Robot Learning.
- 10. Wang, K., Hou, W., Hong, L., & Guo, J. (2025). Smart transparency: A user-centered approach to improving human–machine interaction in high-risk supervisory control tasks. Electronics, 14(3), 420.
- 11. Yang, Y., Wang, Q., Li, C., Hu, H., Wu, C., Jiang, Y., ... & Xu, B. (2025, May). Fewer may be better: Enhancing offline reinforcement learning with reduced dataset. In International Conference on Learning Representations (Vol. 2025, pp. 7076-7098).
- 12. Cheng, Xiaoyuan, et al. "How Does the Lagrangian Guide Safe Reinforcement Learning through Diffusion Models?." arXiv preprint arXiv:2602.02924 (2026).
- 13. Wan, H., Cheng, J., Deng, Y., Wu, D., Chen, Y., Lin, Z., ... & Ji, X. Towards Physics Aware Embodied Control with Graph based Object-centric Learning. ACM Transactions on Cyber-Physical Systems.
- 14. Xiao, Q., & Ghaffari, M. (2026). LaWM: Least Action World Models for Long-Horizon Physical Consistency from Visual Observations. arXiv preprint arXiv:2605.08279.
- 15. Hu, Y., Xi, A., Xiao, Q., Isaacson, S., Liu, H. X., Vasudevan, R., & Ghaffari, M. (2026). LongNav-R1: Horizon-Adaptive Multi-Turn RL for Long-Horizon VLA Navigation. arXiv preprint arXiv:2602.12351.
- 16. Li, Q., Luo, H., Cheng, H., Deng, Y., Sun, W., Li, W., & Liu, Z. (2023). Incipient fault detection in power distribution system: A time-frequency embedded deep-learning-based approach. IEEE Transactions on Instrumentation and Measurement, 72, 1–14.
- 17. Wang, C., Tan, C., Zhu, L., Yang, R., & Hong, J. (2024, September). Neural Active Sensing Vision and Manipulation for Cooperative Agents. In 2024 5th International Conference on Big Data & Artificial Intelligence & Software Engineering (ICBASE) (pp. 942-947). IEEE.
- 18. Wan, H., Zhang, Y., Wang, J., Wu, D., Li, M., Chen, X., ... & Ji, X. (2025). Toward universal embodied planning in scalable heterogeneous field robots collaboration and control. Journal of Field Robotics, 42(5), 2318-2336.
- 19. Shi, T., ElSamadisy, O., & Abdulhai, B. (2026). CoopSECRM2D: A Safe, Efficient, and Comfortable Multi‐Agent Reinforcement Learning Model for Cooperative On‐Ramp Merging. IET Intelligent Transport Systems, 20 (1), e70256.
- 20. Wang Y, Fang Y, Wang T, et al. Dreamnav: A trajectory-based imaginative framework for zero-shot vision-and-language navigation[J]. arXiv preprint arXiv:2509.11197, 2025.
- 21. Yao, Jingyu, et al. "Explainable AI Enhanced Traffic Forecasting for Proactive Autonomous Vehicle Routing." 2026 IEEE Conference on Technologies for Sustainability (SusTech). IEEE, 2026.
- 22. Mnih, V., Badia, A. P., Mirza, M., Graves, A., Harley, T., Lillicrap, T., et al. (2016). Asynchronous methods for deep reinforcement learning. In Proceedings of ICML.
- 23. Levine, S., Pastor, P., Krizhevsky, A., & Quillen, D. (2018). Learning hand-eye coordination for robotic grasping with deep learning and large-scale data collection. International Journal of Robotics Research, 37(4–5), 421–436.
- 24. Li, X., Wang, P., Li, G., & Zhang, Y. (2023). Design of a novel self-test-on-chip interface ASIC for capacitive accelerometers. IEEE Transactions on Circuits and Systems I: Regular Papers, 70(7), 2834-2843.
- 25. Bojarski, M., Del Testa, D., Dworakowski, D., Firner, B., Flepp, B., Goyal, P., et al. (2016). End to end learning for self-driving cars. arXiv preprint arXiv:1604.07316.
- 26. Levine, S., & Koltun, V. (2013). Guided policy search. In Proceedings of ICML.
- 27. Luo, Y., Chen, Y., Duan, Z., He, Y., Tang, H., & Pan, C. (2026). Fault Tolerant Control for Heterogeneous UAVs-UGVs Coordinative System with Adaptive Safety Switching based on DMPC. IEEE Transactions on Aerospace and Electronic Systems.
- 28. Prescribed-time tracking control of MIMO nonlinear systems with nonvanishing uncertainties
- 29. Luo, P., & Ding, X. (2026). C-MAPPO-TSC: A Collaborative Multi-Agent Reinforcement Learning Framework for Urban Traffic Signal Control and Congestion Mitigation. Fundamental Scientific Reports in Multidisciplinary Areas, 2 (02), 170-181.
- 30. Schulman, J., Wolski, F., Dhariwal, P., Radford, A., & Klimov, O. (2017). Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347.
- 31. Zhang, Luyan. "MCP: A Control-Theoretic Orchestration Framework for Synergistic Efficiency and Interpretability in Multimodal Large Language Models." arXiv preprint arXiv:2509.16597 (2025).
- 32. Wang Y, Fang Y, Wang T, et al. What Limits Vision-and-Language Navigation?[J]. arXiv preprint arXiv:2605.13328, 2026.
- 33. Mur-Artal, R., & Tardós, J. D. (2017). ORB-SLAM2: An open-source SLAM system for monocular, stereo, and RGB-D cameras. IEEE Transactions on Robotics, 33(5), 1255–1262.