[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-seo-140840-105":3,"detail-sidebar-cat-0-en-105":75,"doc-detail-140840-en":125},{"code":4,"msg":5,"data":6},0,"ok",{"site_id":7,"language":8,"slug":9,"title":10,"keywords":11,"description":12,"schema_data":13,"social_meta":68,"head_meta":70,"extra_data":72,"updated_unix":74},105,"en","learning-distilled-collaboration-graph-for-multi-agent-perception-supplemental-material","Learning Distilled Collaboration Graph for Multi-Agent Perception - Supplemental Material","","Supplementary material for Learning Distilled Collaboration Graph for Multi-Agent Perception, detailing dataset construction and model internals. It describes a simulated multi-agent 3D perception dataset targeting vehicle, bicycle, and person detection in LiDAR point clouds, including CARLA vehicle taxonomy and automated 3D box retrieval. It explains CARLA–SUMO traffic-flow co-simulation, log-to-scene generation, and an extended nuScenes-style dataset format. It further specifies the student/teacher encoder-decoder architecture built on MotionNet, including BEV tensor dimensions and decoder layer connections, plus visualization references for bounding boxes and projected point clouds.",{"@graph":14,"@context":67},[15,34,50],{"@type":16,"itemListElement":17},"BreadcrumbList",[18,23,27,31],{"item":19,"name":20,"@type":21,"position":22},"https://docshare.wps.com","Home","ListItem",1,{"item":24,"name":25,"@type":21,"position":26},"https://docshare.wps.com/document/","Document",2,{"item":28,"name":29,"@type":21,"position":30},"https://docshare.wps.com/document/research-report/","Research & Report",3,{"item":32,"name":10,"@type":21,"position":33},"https://docshare.wps.com/document/learning-distilled-collaboration-graph-for-multi-agent-perception-supplemental-material/140840/",4,{"url":32,"name":10,"@type":35,"author":36,"headline":10,"publisher":39,"fileFormat":42,"inLanguage":8,"description":12,"dateModified":43,"datePublished":44,"encodingFormat":42,"isAccessibleForFree":45,"interactionStatistic":46},"DigitalDocument",{"name":37,"@type":38},"Oliver Hayes","Person",{"url":19,"name":40,"@type":41},"DocShare","Organization","application/pdf","2026-09-11","2026-08-25",true,{"@type":47,"interactionType":48,"userInteractionCount":33},"InteractionCounter",{"@type":49},"ViewAction",{"@type":51,"mainEntity":52},"FAQPage",[53,59,63],{"name":54,"@type":55,"acceptedAnswer":56},"What targets and data modality does the dataset focus on?","Question",{"text":57,"@type":58},"The dataset targets vehicle, bicycle, and person detection using 3D point clouds. It reports vehicle detection results while leaving bicycle and person detection as follow-up works.","Answer",{"name":60,"@type":55,"acceptedAnswer":61},"How are traffic scenes generated for multi-agent data recording?",{"text":62,"@type":58},"Traffic flow is simulated with CARLA–SUMO co-simulation, where vehicles are spawned in CARLA via SUMO and managed by the Traffic Manager. Scenes are sampled from recorded logs across intersections, with multiple agents per scene.",{"name":64,"@type":55,"acceptedAnswer":65},"What is the backbone and how is the model’s encoder-decoder structured?",{"text":66,"@type":58},"The model uses MotionNet as the backbone, employing an encoder-decoder architecture with skip connections. The document specifies the BEV input dimension and gives layer-by-layer architectures for the student/teacher encoder and decoder, including intermediate feature flow.","https://schema.org",{"og:url":32,"og:type":69,"og:title":10,"og:site_name":40,"og:description":12},"article",{"robots":71,"canonical":32},"index,follow",{"doc_id":73,"site_id":7},140840,1787644735,{"code":4,"msg":76,"data":77},"success",[78,82,86,90,95,100,105,109,114,117,121],{"id":22,"doc_module":4,"doc_module_name":25,"category_name":79,"show_sort_weight":80,"slug":81},"Story & Novel",90,"story-novel",{"id":26,"doc_module":4,"doc_module_name":25,"category_name":83,"show_sort_weight":84,"slug":85},"Literature",80,"literature",{"id":33,"doc_module":4,"doc_module_name":25,"category_name":87,"show_sort_weight":88,"slug":89},"Exam",70,"exam",{"id":91,"doc_module":4,"doc_module_name":25,"category_name":92,"show_sort_weight":93,"slug":94},5,"Comic",60,"comic",{"id":96,"doc_module":4,"doc_module_name":25,"category_name":97,"show_sort_weight":98,"slug":99},6,"Technology",50,"technology",{"id":101,"doc_module":4,"doc_module_name":25,"category_name":102,"show_sort_weight":103,"slug":104},7,"Healthcare",40,"healthcare",{"id":106,"doc_module":4,"doc_module_name":25,"category_name":29,"show_sort_weight":107,"slug":108},8,30,"research-report",{"id":110,"doc_module":4,"doc_module_name":25,"category_name":111,"show_sort_weight":112,"slug":113},9,"Religion & Spirituality",20,"religion-spirituality",{"id":112,"doc_module":4,"doc_module_name":25,"category_name":115,"show_sort_weight":112,"slug":116},"World Cup","world-cup",{"id":118,"doc_module":4,"doc_module_name":25,"category_name":119,"show_sort_weight":118,"slug":120},10,"Lifestyle","lifestyle",{"id":122,"doc_module":4,"doc_module_name":25,"category_name":123,"show_sort_weight":91,"slug":124},19,"General","general",{"code":4,"msg":76,"data":126},{"doc_id":73,"user_id":127,"nickname":37,"user_avatar":128,"doc_module":4,"category_id":106,"category_name":29,"doc_title":10,"doc_description":12,"doc_content":129,"file_id":130,"file_url":131,"file_type":132,"file_size":133,"view_count":33,"is_deleted":4,"is_public":22,"is_downloadable":22,"audit_status":22,"page_count":91,"language":134,"language_code":8,"site_id":7,"html_lang":8,"table_of_contents":135,"faqs":136,"seo_title":137,"seo_description":12,"update_tm":74,"read_time":138},687207020761,"https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d","Supplementary Material: Learning Distilled Collaboration Graph for Multi-Agent Perception  \nYiming Li  \nNew York University [yimingli@nyu.edu](yimingli@nyu.edu)  \nShunli Ren  \nShanghai Jiao Tong University [renshunli@sjtu.edu.cn](renshunli@sjtu.edu.cn)  \nPengxiang Wu  \nRutgers University [pxiangwu@gmail.com](pxiangwu@gmail.com)  \nSiheng Chen􀀃 Shanghai Jiao Tong University [sihengc@sjtu.edu.cn](sihengc@sjtu.edu.cn)  \nChen Feng􀀃 New York University [cfeng@nyu.edu](cfeng@nyu.edu)  \nWenjun Zhang  \nShanghai Jiao Tong University [zhangwenjun@sjtu.edu.cn](zhangwenjun@sjtu.edu.cn)  \nI Detailed information of the dataset  \nVehicle type and annotation. Our dataset targets vehicle, bicycle and person detection in 3D point cloud, and we report the results of vehicle detection and leave the bicycle and person detection asthe follow-up works. Noted that LiDAR point cloud would not capture car/person identity and thus 3D detection in point cloud does not involve privacy issue. There are twenty-one kinds of cars with various sizes and shapes in our simulated dataset, the names of vehicles in CARLA are listed below.  \n[vehicle.bmw.grandtourer](vehicle.bmw.grandtourer) [vehicle.bmw.isetta](vehicle.bmw.isetta) vehicle.chevrolet.impala [vehicle.nissan.patrol](vehicle.nissan.patrol) vehicle.tesla.cybertruck vehicle.tesla.model3 [vehicle.mini.cooperst](vehicle.mini.cooperst) [vehicle.volkswagen.t2](vehicle.volkswagen.t2) [vehicle.toyota.prius](vehicle.toyota.prius)[ ](vehicle.toyota.prius)vehicle.citroen.c3 [vehicle.dodge_charger.police](vehicle.dodge_charger.police) [vehicle.audi.tt](vehicle.audi.tt)[ ](vehicle.audi.tt)vehicle.mustang.mustang [vehicle.nissan.micra](vehicle.nissan.micra) [vehicle.audi.a2](vehicle.audi.a2)[ ](vehicle.audi.a2)[vehicle.jeep.wrangler_rubicon](vehicle.jeep.wrangler_rubicon) vehicle.carlamotors.carlacola [vehicle.audi.etron](vehicle.audi.etron) vehicle.mercedes-benz.coupe [vehicle.lincoln.mkz2017](vehicle.lincoln.mkz2017) [vehicle.seat.leon](vehicle.seat.leon)  \nThe 3D bounding boxes of different vehicles can be readily obtained without human annotations, and the LiDAR point cloud is aligned well with the camera image, as shown in Fig. I.  \nCARLA-SUMO co-simulation. We use CARLA-SUMO co-simulation for trafﬁc ﬂow simulation and data recording. Vehicles are spawned in CARLA via SUMO, and managed by the Trafﬁc Manager. The script spawn_npc_sumo:py provided by CARLA can automatically generate a SUMO network in a certain town, and can produce random routes and make the vehicles roam around, seen in Fig. II. Five hundred vehicles are spawned in Town05 and we record a log ﬁle with a length of ﬁve minutes, then we read out one hundred scenes from the log ﬁle at different intersections. Each scene includes a duration of twenty seconds, and there are totally M (M = 2 ; 3 ; 4 ; 5) agents in a scene. Several examples of the generated scenes are shown in Fig. III.  \nDataset format. We employ the dataset format of the nuScenes and extend it to multi-agent scenarios, seen in Fig. IV. Each log ﬁle can produce 100 scenes, and each scene includes 100 frames. Each frame covers multiple samples generated from multiple agents at the same timestamp. A sample includes the ego-pose of the agent, the sensor calibration information, and the corresponding annotations of its surrounding vehicles. Given a recorded log ﬁle, the dataset based on the log ﬁle can be generated  \n􀀃 Corresponding authors.  \n35th Conference on Neural Information Processing Systems (NeurIPS 2021) .  \nautomatically with our tool, which does not require laborious manual annotations. Note that our dataset can be further enlarged to boost the object categories and trafﬁc scenarios.  \nII Detailed architecture of the model  \nWe use the main architecture of MotionNet [32] as our backbone, which uses an encoder-decoder architecture with skip connection. The input BEV map's dimension is (c; w; h) = (13; 256 ; 256) .  \nII.1 Architecture of student/teacher encoder  \nWe describe the arc","cbCaiiAaCyF1YhXc","https://ap.wps.com/l/cbCaiiAaCyF1YhXc","pdf",2124216,"English","# Detailed information of the dataset\n## Vehicle type and annotation\n## CARLA-SUMO co-simulation\n## Dataset format\n# Detailed architecture of the model\n## Architecture of student/teacher encoder\n## Architecture of student/teacher decoder","[{\"question\":\"What targets and data modality does the dataset focus on?\",\"answer\":\"The dataset targets vehicle, bicycle, and person detection using 3D point clouds. It reports vehicle detection results while leaving bicycle and person detection as follow-up works.\"},{\"question\":\"How are traffic scenes generated for multi-agent data recording?\",\"answer\":\"Traffic flow is simulated with CARLA–SUMO co-simulation, where vehicles are spawned in CARLA via SUMO and managed by the Traffic Manager. Scenes are sampled from recorded logs across intersections, with multiple agents per scene.\"},{\"question\":\"What is the backbone and how is the model’s encoder-decoder structured?\",\"answer\":\"The model uses MotionNet as the backbone, employing an encoder-decoder architecture with skip connections. The document specifies the BEV input dimension and gives layer-by-layer architectures for the student/teacher encoder and decoder, including intermediate feature flow.\"}]","Learning Distilled Collaboration Graph for Multi-Agent Perception - Supplemental Material | PDF",13]