[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86308-en":3,"doc-seo-86308-105":30,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86308,13056703020460,"Valentina","https://ap-avatar.wpscdn.com/avatar/be000253dac470eee5d?_k=1778207105932848923",8,"Research & Report","JobHop v2: A Large-Scale Career Trajectory Dataset from Unstructured Resumes","Large-scale, richly annotated career-trajectory data enables workforce planning, job recommendation, and labour market analysis, yet existing public resources remain limited in size, reuse permissions, or reliance on pre-standardized occupational codes. JobHop v2 improves the JobHop benchmark using end-to-end LLM extraction from ~440,000 pseudonymized multilingual resumes, released as 355,315 trajectories. Each trajectory is annotated with ESCO occupational codes, quarter-level temporal signals, and normalized five-level education attainment. The redesigned reasoning-controlled pipeline with retry achieves a 100% JSON parse rate and highest extraction quality versus multiple annotation baselines.","JobHop v2: A Large-Scale Career Trajectory Dataset from Unstructured Resumes  \nIman Johary* , Guillaume Bied , Alexandru C. Mara and Tijl De Bie  \nAIDA-IDLab, Department of Electronics and Information Systems, Ghent University, Ghent, Belgium  \nAbstract  \nLarge-scale, richly annotated career trajectory data underpins workforce planning, job recommendation, and labour market analysis, yet publicly available datasets are either small, closed to independent use, or built from pre-standardized occupational codes with LLM-synthesized rather than authentic free text. We present JobHop v2, an improved version of the publicly available JobHop dataset, constructed through end-to-end large language model (LLM) extraction from a corpus of ∼440 ,000 pseudonymized, multilingual resumes provided by VDAB, the Flemish Public Employment Service. The released dataset comprises 355 ,315 career trajectories annotated with ESCO occupational codes, quarter-level temporal information, and normalized five-level education attainment, broadening both the coverage and the annotation richness of the original release. Relative to v1, JobHop v2 introduces a redesigned extraction pipeline based on reasoning-controlled LLM inference with a retry mechanism (achieving a 100% JSON parse rate), a richer extraction schema, and a revised evaluation protocol scored against three complementary annotation baselines. Evaluated against these baselines, our best extractor comes closest to the inter-annotator agreement ceiling among all compared models, trailing it by only  \n1.1–2.7 percentage points. The dataset and code are publicly released to support reproducible career-trajectory research.  \nKeywords  \ncareer trajectories, dataset, ESCO, information extraction, labour market analysis, large language models, job recommendation  \n1. Introduction  \nUnderstanding how careers evolve (which roles people transition into, when, and from what educational background) has broad practical value: it supports evidence-based labour market policy, powers job recommendation systems, and enables career counseling at scale [1, 2] . The limiting factor for computational career analysis is data. Resumes are the richest naturally occurring source of career-trajectory information: unlike job postings, which capture only open roles, or administrative records, which rarely include job details, resumes document the full arc of a working life (job titles, responsibilities, education, and the timing of transitions) in a single document. Yet their unstructured, multilingual, and heterogeneous nature has long prevented large-scale systematic use.  \nRecent large language models (LLMs) have changed this calculus. By combining open-ended text comprehension with structured reasoning, LLMs can normalize job titles to occupational taxonomies, resolve temporal ambiguities, and extract skill inventories from free-form descriptions ata scale and accuracy that rule-based and sequence-labeling pipelines could not approach [3, 4] . It is now possible to source rich career trajectory data containing temporal and educational signals from unstructured resumes, enabling large resume corpora to be used for career-path analysis and modeling.  \nDespite this opportunity, the field still lacks a suitable public benchmark. Existing public datasets are either small (Decorte et al. [1] release 2 , 164 career histories and OpenResume [5] only 301 real resumes), built from pre-standardized occupational codes rather than raw text [2], or derived from proprietary platforms [6, 7, 8] that are not publicly available for independent use. JobHop [4] was, to our knowledge,  \nRecSys in HR 2026: The 6th Workshop on Recommender Systems for Human Resources, in conjunction with the 20th ACM Conference on Recommender Systems, September 2026, Prague, Czech Republic  \n* Corresponding author.  \n$ [iman.johary@ugent.be](iman.johary@ugent.be) (I. Johary); [guillaume.bied@ugent.be](guillaume.bied@ugent.be)[ ](guillaume.bied@ugent.be)(G. Bi","cbCaikldqkKm8gxB","https://ap.wps.com/l/cbCaikldqkKm8gxB","pdf",966185,3,1,9,"English","en",105,"# Introduction\n## Motivation and problem setting\n## Why resumes and why LLMs\n## Limitations of existing datasets\n# JobHop v2 overview\n## Data source and coverage\n## Annotation scope (ESCO, time, education)\n## Extraction pipeline improvements\n## Evaluation against baselines","[{\"question\":\"What annotations and output quality improvements does JobHop v2 introduce?\",\"answer\":\"The dataset includes ESCO occupational codes, quarter-level temporal information, and normalized five-level education attainment. A redesigned reasoning-controlled extraction pipeline with a multi-step retry mechanism achieves a 100% JSON parse rate and improved accuracy measured against multiple annotation baselines.\"}]",1784210377,23,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":28},"jobhop-v2-a-large-scale-career-trajectory-dataset-from-unstructured-resumes","",{"@graph":36,"@context":77},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/jobhop-v2-a-large-scale-career-trajectory-dataset-from-unstructured-resumes/86308/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"What annotations and output quality improvements does JobHop v2 introduce?","Question",{"text":75,"@type":76},"The dataset includes ESCO occupational codes, quarter-level temporal information, and normalized five-level education attainment. A redesigned reasoning-controlled extraction pipeline with a multi-step retry mechanism achieves a 100% JSON parse rate and improved accuracy measured against multiple annotation baselines.","Answer","https://schema.org",{"og:url":51,"og:type":79,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":81,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":84},[85,89,93,97,102,107,112,115,119,122,126],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":46,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":46,"category_name":109,"show_sort_weight":110,"slug":111},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":113,"slug":114},30,"research-report",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},"Religion & Spirituality",20,"religion-spirituality",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":120,"show_sort_weight":117,"slug":121},"World Cup","world-cup",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":123,"slug":125},10,"Lifestyle","lifestyle",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":98,"slug":129},19,"General","general"]