{"id":89,"date":"2024-02-22T22:38:09","date_gmt":"2024-02-22T22:38:09","guid":{"rendered":"https:\/\/mrsdprojects.ri.cmu.edu\/2024teamg\/?page_id=89"},"modified":"2024-12-11T06:07:45","modified_gmt":"2024-12-11T06:07:45","slug":"system-implementation","status":"publish","type":"page","link":"https:\/\/mrsdprojects.ri.cmu.edu\/2024teamg\/system\/system-implementation\/","title":{"rendered":"System Implementation"},"content":{"rendered":"\n<hr class=\"wp-block-separator has-alpha-channel-opacity\" \/>\n\n\n\n<h2 class=\"wp-block-heading has-text-align-center\"><strong>Perception Subsystem<\/strong><\/h2>\n\n\n\n<div class=\"wp-block-media-text is-stacked-on-mobile\" style=\"grid-template-columns:47% auto\"><figure class=\"wp-block-media-text__media\"><img loading=\"lazy\" decoding=\"async\" width=\"456\" height=\"788\" src=\"https:\/\/mrsdprojects.ri.cmu.edu\/2024teamg\/wp-content\/uploads\/sites\/80\/2024\/04\/\u622a\u5c4f2024-04-30-17.03.09.png\" alt=\"\" class=\"wp-image-182 size-full\" srcset=\"https:\/\/mrsdprojects.ri.cmu.edu\/2024teamg\/wp-content\/uploads\/sites\/80\/2024\/04\/\u622a\u5c4f2024-04-30-17.03.09.png 456w, https:\/\/mrsdprojects.ri.cmu.edu\/2024teamg\/wp-content\/uploads\/sites\/80\/2024\/04\/\u622a\u5c4f2024-04-30-17.03.09-174x300.png 174w, https:\/\/mrsdprojects.ri.cmu.edu\/2024teamg\/wp-content\/uploads\/sites\/80\/2024\/04\/\u622a\u5c4f2024-04-30-17.03.09-360x622.png 360w, https:\/\/mrsdprojects.ri.cmu.edu\/2024teamg\/wp-content\/uploads\/sites\/80\/2024\/04\/\u622a\u5c4f2024-04-30-17.03.09-250x432.png 250w, https:\/\/mrsdprojects.ri.cmu.edu\/2024teamg\/wp-content\/uploads\/sites\/80\/2024\/04\/\u622a\u5c4f2024-04-30-17.03.09-100x173.png 100w\" sizes=\"auto, (max-width: 456px) 100vw, 456px\" \/><\/figure><div class=\"wp-block-media-text__content\">\n<ul class=\"wp-block-list\">\n<li><strong>Human Pose Detection Module<\/strong>\n<ul class=\"wp-block-list\">\n<li>The current Human Pose Detection Model we use is RTMPose, which takes the streaming images from the drone and outputs 2D pose (17&#215;2 key points)<\/li>\n\n\n\n<li>Then, we use TransformerV2 to leverage these 17 2D key points into 3D key points.<\/li>\n<\/ul>\n<\/li>\n\n\n\n<li><strong>Human 3D Pose to LBS Motion Module<\/strong>\n<ul class=\"wp-block-list\">\n<li>An RNN network is currently trained using the dataset from Audio2Photoreal.<\/li>\n\n\n\n<li>The model feed takes the 3D Pose(17&#215;3 Keypoints) and output LBS motion vector(104&#215;1)<\/li>\n<\/ul>\n<\/li>\n\n\n\n<li><strong>Drone Depth Estimation Module<\/strong>\n<ul class=\"wp-block-list\">\n<li>We first use <strong>PoseFormer V2<\/strong> and <strong>Zoe Depth<\/strong> to get the relative depth map and then estimate the distance between the human and the drone camera.<\/li>\n<\/ul>\n<\/li>\n<\/ul>\n<\/div><\/div>\n\n\n\n<hr class=\"wp-block-separator alignwide has-alpha-channel-opacity\" \/>\n\n\n\n<h2 class=\"wp-block-heading has-text-align-center\">                                                                                   Human 3D Avatar Subsystem<\/h2>\n\n\n\n<div class=\"wp-block-media-text has-media-on-the-right is-stacked-on-mobile\"><div class=\"wp-block-media-text__content\">\n<ul class=\"wp-block-list\">\n<li><strong>Decoder<\/strong>\n<ul class=\"wp-block-list\">\n<li>The decoder we used is provided by our sponsor, Meta. <\/li>\n\n\n\n<li>It renders the texture of meta codec avatars given lbs motion.<\/li>\n\n\n\n<li>The decoder itself is around 1-2 FPS.<\/li>\n<\/ul>\n<\/li>\n\n\n\n<li><strong>Renderer<\/strong>\n<ul class=\"wp-block-list\">\n<li>From the output of the Decoder, we could use the Renderer to generate images of the avatar given the camera pose.<\/li>\n\n\n\n<li>To achieve 30 FPS visualization, we downgrade the resolution of the generated avatar by 4x times, achieving 4 FPS for the decoder. <\/li>\n<\/ul>\n<\/li>\n\n\n\n<li><strong>Interpolation Module<\/strong>\n<ul class=\"wp-block-list\">\n<li>Due to the decoder&#8217;s inefficiency, we ran it every 8th frame and generated the frames in between using the interpolation method.<\/li>\n\n\n\n<li>With the interpolation, we could eventually achieve around 32 FPS. <\/li>\n<\/ul>\n<\/li>\n<\/ul>\n<\/div><figure class=\"wp-block-media-text__media\"><img loading=\"lazy\" decoding=\"async\" width=\"762\" height=\"1024\" src=\"https:\/\/mrsdprojects.ri.cmu.edu\/2024teamg\/wp-content\/uploads\/sites\/80\/2024\/04\/avatar-762x1024.gif\" alt=\"\" class=\"wp-image-173 size-full\" srcset=\"https:\/\/mrsdprojects.ri.cmu.edu\/2024teamg\/wp-content\/uploads\/sites\/80\/2024\/04\/avatar-762x1024.gif 762w, https:\/\/mrsdprojects.ri.cmu.edu\/2024teamg\/wp-content\/uploads\/sites\/80\/2024\/04\/avatar-223x300.gif 223w, https:\/\/mrsdprojects.ri.cmu.edu\/2024teamg\/wp-content\/uploads\/sites\/80\/2024\/04\/avatar-768x1032.gif 768w, https:\/\/mrsdprojects.ri.cmu.edu\/2024teamg\/wp-content\/uploads\/sites\/80\/2024\/04\/avatar-700x940.gif 700w, https:\/\/mrsdprojects.ri.cmu.edu\/2024teamg\/wp-content\/uploads\/sites\/80\/2024\/04\/avatar-520x699.gif 520w, https:\/\/mrsdprojects.ri.cmu.edu\/2024teamg\/wp-content\/uploads\/sites\/80\/2024\/04\/avatar-360x484.gif 360w, https:\/\/mrsdprojects.ri.cmu.edu\/2024teamg\/wp-content\/uploads\/sites\/80\/2024\/04\/avatar-250x336.gif 250w, https:\/\/mrsdprojects.ri.cmu.edu\/2024teamg\/wp-content\/uploads\/sites\/80\/2024\/04\/avatar-100x134.gif 100w\" sizes=\"auto, (max-width: 762px) 100vw, 762px\" \/><\/figure><\/div>\n\n\n\n<hr class=\"wp-block-separator alignwide has-alpha-channel-opacity\" \/>\n\n\n\n<h2 class=\"wp-block-heading has-text-align-center\">System Communication Subsystem<\/h2>\n\n\n\n<div class=\"wp-block-media-text is-stacked-on-mobile\"><figure class=\"wp-block-media-text__media\"><img loading=\"lazy\" decoding=\"async\" width=\"393\" height=\"540\" src=\"https:\/\/mrsdprojects.ri.cmu.edu\/2024teamg\/wp-content\/uploads\/sites\/80\/2024\/04\/drone.png\" alt=\"\" class=\"wp-image-180 size-full\" srcset=\"https:\/\/mrsdprojects.ri.cmu.edu\/2024teamg\/wp-content\/uploads\/sites\/80\/2024\/04\/drone.png 393w, https:\/\/mrsdprojects.ri.cmu.edu\/2024teamg\/wp-content\/uploads\/sites\/80\/2024\/04\/drone-218x300.png 218w, https:\/\/mrsdprojects.ri.cmu.edu\/2024teamg\/wp-content\/uploads\/sites\/80\/2024\/04\/drone-360x495.png 360w, https:\/\/mrsdprojects.ri.cmu.edu\/2024teamg\/wp-content\/uploads\/sites\/80\/2024\/04\/drone-250x344.png 250w, https:\/\/mrsdprojects.ri.cmu.edu\/2024teamg\/wp-content\/uploads\/sites\/80\/2024\/04\/drone-100x137.png 100w\" sizes=\"auto, (max-width: 393px) 100vw, 393px\" \/><\/figure><div class=\"wp-block-media-text__content\">\n<ul class=\"wp-block-list\">\n<li><strong>ROS System for Perception Module<\/strong>\n<ul class=\"wp-block-list\">\n<li>We are using ROS nodes between each module to communicate and convey messages. Currently, we have a Video Capture node(Capture Drone Image through HD Capture Card from Remote Controller), a 2DPose Node(Detect 2D Pose and predict the LBS Notion), a Decoder Node(take the LBS motion and render avatar), and an Interpolation Node(take the rendered images and interpolate frames between them).<\/li>\n<\/ul>\n<\/li>\n\n\n\n<li><strong>Drone Control Module<\/strong>\n<ul class=\"wp-block-list\">\n<li>We can control the drone through Payload SDK using an e-port and an onboard computer.<\/li>\n\n\n\n<li>We have developed key points trajectory following and autonomously taking off, hovering, and landing.<\/li>\n<\/ul>\n<\/li>\n<\/ul>\n<\/div><\/div>\n\n\n\n<hr class=\"wp-block-separator alignwide has-alpha-channel-opacity\" \/>\n\n\n\n<h2 class=\"wp-block-heading has-text-align-center\">Human Following Subsystem<\/h2>\n\n\n\n<div class=\"wp-block-media-text has-media-on-the-right is-stacked-on-mobile\"><div class=\"wp-block-media-text__content\">\n<ul class=\"wp-block-list\">\n<li>Yaw Maneuver with Head Pose: \n<ul class=\"wp-block-list\">\n<li>Leveraged head pose estimation derived from 2-D human pose detection to execute yaw maneuvers, enabling the drone to adjust its orientation to maintain visual tracking of the human subject.<\/li>\n<\/ul>\n<\/li>\n\n\n\n<li>PID Controller for Pose Centering: \n<ul class=\"wp-block-list\">\n<li>Developed and implemented a Proportional-IntegralDerivative (PID) controller to maintain the human subject\u2019s pose centered within the camera frame. While the PID controller operates smoothly within a speed range of 0.1 to 0.4 radians per second, it exhibits instability and jittering behavior beyond this range, necessitating further optimization for enhanced stability across variable speeds.<\/li>\n<\/ul>\n<\/li>\n<\/ul>\n<\/div><figure class=\"wp-block-media-text__media\"><img loading=\"lazy\" decoding=\"async\" width=\"680\" height=\"501\" src=\"https:\/\/mrsdprojects.ri.cmu.edu\/2024teamg\/wp-content\/uploads\/sites\/80\/2024\/02\/\u622a\u5c4f2024-02-23-20.31.42.png\" alt=\"\" class=\"wp-image-125 size-full\" srcset=\"https:\/\/mrsdprojects.ri.cmu.edu\/2024teamg\/wp-content\/uploads\/sites\/80\/2024\/02\/\u622a\u5c4f2024-02-23-20.31.42.png 680w, https:\/\/mrsdprojects.ri.cmu.edu\/2024teamg\/wp-content\/uploads\/sites\/80\/2024\/02\/\u622a\u5c4f2024-02-23-20.31.42-300x221.png 300w\" sizes=\"auto, (max-width: 680px) 100vw, 680px\" \/><\/figure><\/div>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\" \/>\n\n\n\n<h2 class=\"wp-block-heading has-text-align-center\">VR Headset Subsystem<\/h2>\n\n\n\n<p><strong>Windows ROS Receiver<\/strong><br>After encoding the human pose to the LBS Keypints, the set of 104 keypoints are streamed<br>over the ROS 2 network from the Ubuntu System to the Windows system. This is essential due<br>to the fact that the Quest Link drivers are only released by Meta for Windows systems. Another<br>issue that arises due to a Windows system is used is that two different environments that are<br>required to implement the pipeline. This is caused by ROS 2 and Pytorch being incompatible<br>on the same environment.<br><strong>VUER Renderer<\/strong><br>This subsystem receives the 104 keypoints from the ROS 2 environment and renders it<br>using VUER renderer. The VUER Renderer was chosen after initially attempting to render<br>directly using Meta\u2019s DRTK renderer. However due to difficulty in estimating the right transforms for each eye and overall latency in the pipeline, the VUER Renderer was used instead.<br>After the generated 3D mesh is pushed on the VUER interface, we can then stream it to VR Headset using quest link.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Perception Subsystem Human 3D Avatar Subsystem System Communication Subsystem Human [&hellip;]<\/p>\n","protected":false},"author":373,"featured_media":0,"parent":71,"menu_order":0,"comment_status":"closed","ping_status":"closed","template":"page-templates\/page_fullwidth.php","meta":{"footnotes":""},"class_list":["post-89","page","type-page","status-publish","hentry","clearfix"],"_links":{"self":[{"href":"https:\/\/mrsdprojects.ri.cmu.edu\/2024teamg\/wp-json\/wp\/v2\/pages\/89","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/mrsdprojects.ri.cmu.edu\/2024teamg\/wp-json\/wp\/v2\/pages"}],"about":[{"href":"https:\/\/mrsdprojects.ri.cmu.edu\/2024teamg\/wp-json\/wp\/v2\/types\/page"}],"author":[{"embeddable":true,"href":"https:\/\/mrsdprojects.ri.cmu.edu\/2024teamg\/wp-json\/wp\/v2\/users\/373"}],"replies":[{"embeddable":true,"href":"https:\/\/mrsdprojects.ri.cmu.edu\/2024teamg\/wp-json\/wp\/v2\/comments?post=89"}],"version-history":[{"count":13,"href":"https:\/\/mrsdprojects.ri.cmu.edu\/2024teamg\/wp-json\/wp\/v2\/pages\/89\/revisions"}],"predecessor-version":[{"id":239,"href":"https:\/\/mrsdprojects.ri.cmu.edu\/2024teamg\/wp-json\/wp\/v2\/pages\/89\/revisions\/239"}],"up":[{"embeddable":true,"href":"https:\/\/mrsdprojects.ri.cmu.edu\/2024teamg\/wp-json\/wp\/v2\/pages\/71"}],"wp:attachment":[{"href":"https:\/\/mrsdprojects.ri.cmu.edu\/2024teamg\/wp-json\/wp\/v2\/media?parent=89"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}