[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-151599-en":3,"doc-seo-151599-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},151599,962085662650,"Dozel","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","JINX：UNLIMITED LLMS FOR PROBING ALIGNMENT FAILURES","Jinx introduces a helpful-only variant of open-weight LLMs designed to respond to all queries without refusals or safety filtering. Trained without safety-alignment constraints, it preserves the base model’s reasoning and instruction-following abilities while enabling systematic probing of alignment failures and safety boundaries. The paper motivates the need for accessible unconstrained models, then evaluates Jinx across safety, instruction following, general reasoning, and mathematical reasoning, comparing results against original base models and discussing research applications such as data synthesis, red teaming, interpretability, and multi-agent studies.","arXiv :2508 .08243v3 [ cs .CL] 24 Aug 2025  \nJINX: UNLIMITED LLMS FOR PROBING ALIGNMENT FAILURES  \nJiahao Zhao Liwei Dong  \n\" This paper contains text that might be offensive. \"  \nABSTRACT  \nUnlimited, or so-called helpful-only language models are trained without safety alignment constraints and never refuse user queries. They are widely used by leading AI companies as internal tools for red teaming and alignment evaluation. For example, if a safety-aligned model produces harmful outputs similar to an unlimited model, this indicates alignment failures that require further attention. Despite their essential role in assessing alignment, such models are not available to the research community. We introduce Jinx 1 , a helpful-only variant of popular open-weight LLMs. Jinx responds to all queries without refusals or safety filtering, while preserving the base model’s capabilities in reasoning and instruction following. It provides researchers with an accessible tool for probing alignment failures, evaluating safety boundaries, and systematically studying failure modes in language model safety.  \n~~ Safety  ~~Instruction-following~~  ~~  \n~~  ~~General-reasoning~~   ~~Math-reasoning~~  ~~  \n100  \n80  \n60  \n40  \n20  \n0  \n93.75 94.15 93.1  \n\n|  |\n| --- |\n|  |\n|  |\n\nGPQA[main] livemathbench  \n Jinx-Qwen3-235B-A22B-Thinking-2507  Qwen3-235B-A22B-Thinking-2507  \n Jinx-DeepSeek-R1-0528  DeepSeek-R1-0528  \n Jinx-gpt-oss-20b  gpt-oss-20b  \nFigure 1: Jinx main results.  \n1[https://huggingface.co/Jinx-org](https://huggingface.co/Jinx-org)  \n「祸兮福之所倚，福兮祸之所伏。」  \n—–《道德经》  \n1 Introduction  \nThroughout the trajectory of technological advancement, societies have consistently prioritized assessing and mitigating risks associated with emerging technologies [1, 2, 3] . From the early stages of AI development, leading AI companies have deeply embedded safety risk assessment and governance frameworks into their model design and iteration processes [4, 5, 6] . Anthropic’s AI Safety Level (ASL) framework [4] establishes escalating safety, security, and operational standards that correspond to each model’s potential for catastrophic risk. Similarly, OpenAI’s Preparedness Team [5] focuses on tracking, evaluating, and protecting against emerging risks from frontier AI models. This safetyby-design methodology, which embeds protective measures from the conceptual stage and evolves iteratively alongside capability advancement, constitutes a multi-layered defense architecture. These frameworks establish pathways for the industry to address potential catastrophic risks while embodying human-centered AI [7] development philosophy.  \nIn the meantime, academic researchers are actively investigating AI model safety and interpretability, revealing the limitations of existing safety mechanisms. This research primarily assesses the safety alignment mechanisms through three directions: jailbreak attacks [8, 9] use carefully crafted inputs to bypass safety protections and induce harmful content generation; adversarial fine-tuning [10] demonstrates that safety-aligned models may exhibit inappropriate behavioral drift during specific fine-tuning processes; and model interpretability [11, 12] analysis identifies security vulnerabilities and potential failure modes by parsing internal model mechanisms. These research efforts collectively demonstrate that despite current AI systems employing multiple safety alignment strategies, risks of malicious misuse or accidental failure persist.  \nAs LLM scales expand and training processes become more complex, safety alignment itself becomes increasingly challenging. The risk of reward hacking [13] in reinforcement learning-based post-training is growing substantially. To investigate these challenges, Anthropic [14] has explored deceptive alignment phenomena in helpful-only models, revealing risks where models may appear to perform well on the surface while harboring problematic internal behaviors. Similarly, OpenAI’s related research [1","cbCaisC9PsPl1QnW","https://ap.wps.com/l/cbCaisC9PsPl1QnW","pdf",543290,1,10,"English","en",105,"# Introduction\n## Safety risk assessment and governance in industry\n## Academic directions for alignment and safety evaluation\n## Deceptive alignment and helpful-only model research gap\n## Introducing Jinx and its research use cases\n# Empirical Results\n## Evaluation dimensions and comparisons\n## Jinx model derivations and architectures","[{\"question\":\"What problem does the paper address with helpful-only models?\",\"answer\":\"It addresses the lack of accessible helpful-only models for the broader research community, which limits external validation of alignment failures and safety boundaries.\"},{\"question\":\"How does Jinx differ from safety-aligned models?\",\"answer\":\"Jinx is trained without safety-alignment constraints and responds to risk-related queries with near-zero refusal rate, while keeping reasoning and instruction-following capabilities from the base model.\"},{\"question\":\"What are the main research applications of Jinx described in the paper?\",\"answer\":\"The paper outlines data synthesis for improving safety detection robustness, red teaming for assessing deceptive alignment, model interpretability for observing unconstrained behavior, and use in multi-agent systems as a critic or non-cooperative agent.\"}]","JINX：UNLIMITED LLMS FOR PROBING ALIGNMENT FAILURES | PDF",1787844561,25,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"jinx-unlimited-llms-for-probing-alignment-failures","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/jinx-unlimited-llms-for-probing-alignment-failures/151599/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-09-04","2026-08-27",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does the paper address with helpful-only models?","Question",{"text":76,"@type":77},"It addresses the lack of accessible helpful-only models for the broader research community, which limits external validation of alignment failures and safety boundaries.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does Jinx differ from safety-aligned models?",{"text":81,"@type":77},"Jinx is trained without safety-alignment constraints and responds to risk-related queries with near-zero refusal rate, while keeping reasoning and instruction-following capabilities from the base model.",{"name":83,"@type":74,"acceptedAnswer":84},"What are the main research applications of Jinx described in the paper?",{"text":85,"@type":77},"The paper outlines data synthesis for improving safety detection robustness, red teaming for assessing deceptive alignment, model interpretability for observing unconstrained behavior, and use in multi-agent systems as a critic or non-cooperative agent.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":21,"slug":134},"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":107,"slug":138},19,"General","general"]