[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82905-en":3,"doc-seo-82905-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82905,8796095461564,"Liam","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","RepoTrace Browser Assisted Evidence Collection for GitHub Research Datasets","Empirical software engineering studies often assemble datasets from GitHub issues and pull requests, but the browser-to-spreadsheet workflow scatters page evidence, research code, and decision rationale across tabs and files. RepoTrace addresses this by linking GitHub evidence to research interpretation in a local SQLite-backed workspace. The system uses a Chrome side-panel extension, an Express backend, and a React dashboard to capture snapshots, comments, labels, notes, screening choices, review history, and scoped exports. Validation on 20 Matplotlib issues preserves 22 snapshots, 38 comments, 98 annotations, and supports a complete manual collection workflow.","RepoTrace: Browser-Assisted Evidence Collection for GitHub  \nResearch Datasets  \nXue Yao  \nMonash University Melbourne, Australia [xyao0028@student.monash.edu](xyao0028@student.monash.edu)  \nZehua Zhang  \nMonash University Melbourne, Australia [zzha0610@student.monash.edu](zzha0610@student.monash.edu)  \nJiatong Liu  \nMonash University Melbourne, Australia [jliu0455@student.monash.edu](jliu0455@student.monash.edu)  \nYongqiang Tian  \nMonash University Melbourne, Australia [yongqiang.tian@monash.edu](yongqiang.tian@monash.edu)  \narXiv :2607 .05 106v 1 [ cs . SE] 6 Jul 2026  \nAbstract  \nEmpirical software engineering studies frequently build datasets from GitHub issues and pull requests. In many projects, researchers inspect pages in a browser, copy selected fields into spreadsheets, keep side notes in separate documents, and later run scripts to normalize or export the data. This workflow is flexible, but the page evidence, the research codes, and the rationale behind each decision end up spread across tabs and files, which leaves provenance, update tracking, and multi-reviewer labeling hard to audit.  \nRepoTrace is a browser-assisted research tool that collects GitHub issue and pull-request evidence into a local SQLite-backed workspace. It combines a Chrome side-panel extension, an Express backend, and a React dashboard to capture page snapshots, comments, labels, notes, screening and labeling decisions, refresh history, and scoped exports, keeping the source evidence and the research interpretation linked together.  \nA validation pass collected and checked 20 Matplotlib issues across two study projects. The resulting dataset preserves 22 snapshots, 38 comments, 20 research notes, 98 annotations, 20 screening reviews, 20 fix-evidence entries, and 4 simulated unresolved consensus conflicts. The results show that RepoTrace can support a complete local evidence-collection workflow for manually constructed GitHub issue and pull-request datasets.  \nCCS Concepts  \n• Software and its engineering → Software testing and debugging; • Information systems → Web applications.  \nKeywords  \nempirical software engineering, GitHub mining, research datasets, evidence collection, issue tracking  \nACM Reference Format:  \nXue Yao, Zehua Zhang, Jiatong Liu, and Yongqiang Tian. 2026. RepoTrace: Browser-Assisted Evidence Collection for GitHub Research Datasets. In Proceedings of Proceedings of the 35thACM SIGSOFT International Symposium on Software Testing and Analysis (ISSTA’26) . ACM, New York, NY, USA, 5 pages.  \nISSTA’26, Oakland, CA, USA 2026.  \n1 Introduction  \nMany empirical software engineering studies begin the same way: researchers identify relevant GitHub issues or pull requests, inspect their discussion, decide whether each belongs in the study, and assign research labels. In practice this workflow is iterative and evidence-heavy. Researchers read issue bodies, comments, linked pull requests, patches, tests, and maintainer labels, and they must revisit records as those records evolve. They also need to record why a record was included or excluded, which evidence supports each label, and where reviewers disagree. Such needs are common in mining software repositories research, which routinely draws on the evolving socio-technical evidence that GitHub records, such as pull-based development traces and contribution-evaluation signals [6, 9, 11] .  \nTwo common workflows address this task, and each leaves a gap. Spreadsheet-based coding is lightweight and flexible: a row can hold a URL, a few labels, and a note. With enough discipline a researcher can also paste screenshots or archive pages by hand, so the limitation is not that any single field is impossible to store. The difficulty is keeping the page evidence, the evolving research codes, and the record of why a decision was made linked together and queryable as the study progresses, rather than scattered across tabs, files, and ad-hoc columns. At the other extreme, mining scripts collect stru","cbCaiatpo7zwS73v","https://ap.wps.com/l/cbCaiatpo7zwS73v","pdf",554498,4,1,5,"English","en",105,"# Introduction\n## Evidence-Heavy Dataset Construction\n## RepoTrace Workflow and Design\n## Productivity Capabilities","[{\"question\":\"What problem does RepoTrace address in GitHub research dataset construction?\",\"answer\":\"It addresses how page evidence, research labels, and decision rationale get scattered across tabs and files, making provenance, update tracking, and multi-reviewer audits difficult.\"},{\"question\":\"How does RepoTrace keep GitHub evidence linked to research interpretation?\",\"answer\":\"RepoTrace treats each issue or pull request as both a source document and a changing research object, storing browser snapshots, extracted metadata, labels, notes, review decisions, and refresh history together in one local workspace.\"},{\"question\":\"What does the validation study show about RepoTrace’s effectiveness?\",\"answer\":\"A validation pass collected and checked 20 Matplotlib issues, preserving snapshots, comments, notes, annotations, screening reviews, and fix-evidence entries, demonstrating support for a complete local evidence-collection workflow for manually built datasets.\"}]",1784183837,13,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"repotrace-browser-assisted-evidence-collection-for-github-research-datasets","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/repotrace-browser-assisted-evidence-collection-for-github-research-datasets/82905/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does RepoTrace address in GitHub research dataset construction?","Question",{"text":75,"@type":76},"It addresses how page evidence, research labels, and decision rationale get scattered across tabs and files, making provenance, update tracking, and multi-reviewer audits difficult.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does RepoTrace keep GitHub evidence linked to research interpretation?",{"text":80,"@type":76},"RepoTrace treats each issue or pull request as both a source document and a changing research object, storing browser snapshots, extracted metadata, labels, notes, review decisions, and refresh history together in one local workspace.",{"name":82,"@type":73,"acceptedAnswer":83},"What does the validation study show about RepoTrace’s effectiveness?",{"text":84,"@type":76},"A validation pass collected and checked 20 Matplotlib issues, preserving snapshots, comments, notes, annotations, screening reviews, and fix-evidence entries, demonstrating support for a complete local evidence-collection workflow for manually built datasets.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,109,114,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":106,"show_sort_weight":107,"slug":108},"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":22,"slug":137},19,"General","general"]