[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82753-en":3,"doc-seo-82753-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82753,549758252649,"Ivy","https://ap-avatar.wpscdn.com/avatar/8000253669c5317157?_k=1778319167496531819",8,"Research & Report","Responsibility Distribution Estimation in Ego-View Accident Videos with Multimodal Large Language Models","Recent studies on multimodal traffic accident understanding rely mainly on infrastructure-camera footage, satellite imagery, or structured crash records, which are expensive to deploy and maintain and cannot objectively reflect what a driver actually saw before the crash. Ego-view accident videos offer a direct view of the driver’s perspective, enabling reasoning about avoidability and driver responsibility. This paper introduces responsibility distribution estimation, predicting each involved agent’s assigned responsibility percentage via an LLM-assisted labeling pipeline and multimodal model fine-tuning across multiple input settings. Experiments provide a strong benchmark and show that multimodal LLMs can perform this constraint-based reasoning, supporting socially and legally meaningful analysis beyond classification.","Responsibility Distribution Estimation in Ego-View Accident Videos with Multimodal Large Language Models  \nRyosei Tamura  \nKeio University [ryoseitamura@keio.jp](ryoseitamura@keio.jp)  \nAndrew Shin  \nKeio University [shin@inl.ics.keio.ac.jp](shin@inl.ics.keio.ac.jp)  \narXiv :2607 .0359 1v 1 [ cs .CV] 3 Jul 2026  \nAbstract  \nRecent studies on multimodal traffic accident understanding have mainly relied on infrastructure-camera footage, satellite imagery, or structured crash records. However, such data sources are costly to deploy and maintain at large scale, and they cannot objectively capture what the driver was actually able to observe before the accident. In contrast, ego-view accident videos directly represent the driver’s visual perspective, making them suitable for reasoning about avoidability and driver responsibility. In this paper, we introduce responsibility distribution estimation for ego-view traffic accident videos, a new task in which a model predicts the percentage of responsibility assigned to each involved agent. We construct an LLM-assisted responsibility annotation pipeline and fine-tune multimodal large language models under multiple input settings, including raw frames, segmentation-enhanced input, and textual descriptions. Experimental results establish a strong initial benchmark, demonstrating that multimodal LLMs can effectively perform this nuanced, constraint-based reasoning task. Our findings suggest that egocentric accident videos provide a promising foundation for socially and legally meaningful multimodal reasoning beyond conventional accident classification and explanation tasks.  \n1 Introduction  \nRecent advances in large language models (LLMs) and multimodal large language models (MLLMs) have enabled substantial progress in traffic accident analysis, including crash detection, severity prediction, and causal explanation. Existing approaches have primarily relied on structured crash reports, satellite imagery, or infrastructure-camera footage. While these settings are useful for scene-level monitoring, they suffer from two important limitations.  \nFirst, infrastructure-based systems are costly to deploy and maintain at scale. Traffic monitoring  \ncameras and high-resolution sensing infrastructure are typically limited to specific intersections or urban regions, making comprehensive coverage difficult. Second, these external viewpoints cannot objectively represent what the driver was actually able to observe before the accident. As a result, they are less suitable for reasoning about avoidability and driver responsibility from the perspective of the vehicle involved in the accident.  \nIn contrast, ego-view accident videos directly capture the driver’s visual perspective. This perspective enables driver-centered reasoning, such as whether the ego driver had sufficient time to react, whether another agent suddenly entered the scene, and how responsibility should be distributed among the involved participants. Despite this unique property, prior work on ego-centric accident understanding has mainly focused on accident cause explanation and prevention-oriented reasoning, leaving responsibility allocation largely unexplored.  \nIn this work, we propose responsibility distribution estimation for ego-view traffic accident videos. Given an accident video and a set of involved agents, the goal is to predict a responsibility distribution over the agents rather than a single accident label or cause. For example, instead of predicting that a pedestrian alone caused the accident, the model may estimate that the pedestrian is 60% responsible while the ego driver is 40% responsible. This formulation allows models to represent shared responsibility and more closely reflects realworld traffic reasoning. To study this problem, we construct an LLM-assisted responsibility labeling pipeline to generate proportional responsibility distributions. We establish the first benchmark for this task by evaluating multimodal mod","cbCaink4ej5gwqCv","https://ap.wps.com/l/cbCaink4ej5gwqCv","pdf",2081038,3,1,6,"English","en",105,"# Introduction\n## Motivation and limitations of existing data sources\n## Ego-view videos for driver-centered responsibility reasoning\n## Proposed task and approach\n## Contributions\n# Related Work\n## LLMs for traffic crash analysis\n## MLLMs for interpretable accident analysis from imagery","[{\"question\":\"What new task does the paper introduce for ego-view accident videos?\",\"answer\":\"It introduces responsibility distribution estimation, where a model predicts the percentage of responsibility assigned to each involved agent rather than a single label or cause.\"},{\"question\":\"Why are ego-view accident videos useful for reasoning about responsibility?\",\"answer\":\"They directly capture the driver’s visual perspective, enabling driver-centered judgments such as reaction time sufficiency and how responsibility should be shared among participants.\"},{\"question\":\"How do the authors build training data and evaluate multimodal models?\",\"answer\":\"They construct an LLM-assisted responsibility annotation pipeline to generate proportional responsibility distributions, then fine-tune and evaluate multimodal large language models under multiple input conditions including raw frames, segmentation-enhanced inputs, and textual descriptions.\"}]",1784182707,15,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"responsibility-distribution-estimation-in-ego-view-accident-videos-with-multimodal-large-language-models","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/responsibility-distribution-estimation-in-ego-view-accident-videos-with-multimodal-large-language-models/82753/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What new task does the paper introduce for ego-view accident videos?","Question",{"text":75,"@type":76},"It introduces responsibility distribution estimation, where a model predicts the percentage of responsibility assigned to each involved agent rather than a single label or cause.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why are ego-view accident videos useful for reasoning about responsibility?",{"text":80,"@type":76},"They directly capture the driver’s visual perspective, enabling driver-centered judgments such as reaction time sufficiency and how responsibility should be shared among participants.",{"name":82,"@type":73,"acceptedAnswer":83},"How do the authors build training data and evaluate multimodal models?",{"text":84,"@type":76},"They construct an LLM-assisted responsibility annotation pipeline to generate proportional responsibility distributions, then fine-tune and evaluate multimodal large language models under multiple input conditions including raw frames, segmentation-enhanced inputs, and textual descriptions.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]