[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84033-en":3,"doc-seo-84033-105":30,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84033,13056703019404,"Miles","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Beyond the Syntax: Do Security Experts Trust LLMs for NIDS Rule Engineering","Network threats evolve rapidly, turning manual NIDS rule engineering into an operational bottleneck. This preprint investigates human-centered, LLM-based NIDS rule generation by defining a grounded generation framework and validating it via a user study with 10 domain experts. Results show a syntax-semantics paradox: rules are often syntactically correct, yet deployability is limited by low specificity and logic hallucinations in 12% of cases. A SUS score of 67 yields overall skepticism, positioning LLMs as drafting and verification support. Statistical analysis finds large models (≥70B) reliably produce valid syntax, while small models (≤4B) are largely ineffective.","BEYOND THE SYNTAX: DO SECURITY EXPERTS TRUST LLMS  \nFOR NIDS RULE ENGINEERING?  \nA PREPRINT  \nLorenzo Di Filippo  \nSorbonne University  \n4 Pl. Jussieu Paris, France, 75005 [Lorenzo.Di-Filippo@lip6.fr](Lorenzo.Di-Filippo@lip6.fr)  \nEnkeleda Bardhi  \nDelft University of Technology Van Mourik Broekmanweg 5 Delft, The Netherlands, 2628 XE[E.Bardhi-1@tudelft.nl](E.Bardhi-1@tudelft.nl)  \nAndrea Agiollo  \nDelft University of Technology Van Mourik Broekmanweg 5 Delft, The Netherlands, 2628 XE[A.Agiollo-1@tudelft.nl](A.Agiollo-1@tudelft.nl)  \narXiv :2607 .059 16v 1 [ cs .CR] 7 Jul 2026  \nAlessandro Palma  \nSapienza University of Rome Via Ariosto 25 Rome, Italy, 00185 [palma@diag.uniroma1.it](palma@diag.uniroma1.it)  \nSilvia Bonomi  \nSapienza University of Rome Via Ariosto 25  \nRome, Italy, 00185 [bonomi@diag.uniroma1.it](bonomi@diag.uniroma1.it)  \nFernando Kuipers  \nDelft University of Technology Van Mourik Broekmanweg 5 Delft, The Netherlands, 2628 XE[F.A.Kuipers@tudelft.nl](F.A.Kuipers@tudelft.nl)  \nABSTRACT  \nAs network threats evolve, manual NIDS rule engineering has become a critical operational bottleneck.  \nWhile Large Language Models (LLMs) show promise for automating this process, their ability to produce production-ready rules remains unvalidated. This paper presents a human-centered investigation into LLM-based NIDS rule engineering, formalizing a grounded generation framework and evaluating it through a user study with 10 domain experts.  \nOur evaluation reveals a syntax-semantics paradox: although LLMs generate syntactically correct rules, experts find them only partially deployable due to low specificity and logic hallucinations in 12% of cases. While the system received a favorable SUS score of 67, practitioners remain skeptical of its autonomous capabilities, viewing LLMs as support tools for drafting and verification rather than independent generators. Finally, our statistical analysis indicates that while large-scale models ((≥ 70B) consistently produce syntactically valid rules, small models (≤ 4B) are largely ineffective for IDS rule generation.  \nKeywords Intrusion Detection Systems · Rule Engineering · Large Language Models · User Study  \n1 Introduction  \nIn the era of distributed and agentic systems where cyber threats evolve at unprecedented speed, Network Intrusion Detection Systems (NIDSs) serve as the first line of defense, continuously monitoring network traffic to detect malicious activities, thus being the cornerstone of modern network defense [31, 49, 55, 15] . Open-source solutions like Snort 1 and Suricata2 have established as industry standards, distinguished by their robust packet processing capabilities, extensive rule communities, and multi-threaded architecture. These systems rely on signature-based detection engines, where network traffic is analyzed against databases of predefined rules. Each rule comprises a header specifying protocol, IP addresses and ports, alongside an option section encoding detection logic. While this engine delivers high-fidelity detection of known threats, its efficacy remains contingent upon the quality, timeliness, and comprehensive coverage of deployed rulesets. These qualities, in turn, rely on significant human effort to create and maintain such rulesets [27, 23] . This is further exacerbated by the escalating complexity of contemporary threat landscapes, including polymorphic malware, zero-day exploits, and rapidly evolving attack infrastructure. This complexity has transformed  \n1[https://www.snort.org](https://www.snort.org)  \n2[https://suricata.io](https://suricata.io)  \nruleset maintenance into a critical operational bottleneck. As network traffic volumes continue their exponential growth, manual rule engineering emerges not merely as labor-intensive, but fundamentally unscalable [22, 49] . This challenge is well-documented as Security Operation Center (SOC) fatigue, wherein security analysts confront alert fatigue from high false positive rates, syntactic compl","cbCaihqyCfnOQ6S8","https://ap.wps.com/l/cbCaihqyCfnOQ6S8","pdf",1456322,3,1,28,"English","en",105,"# Introduction\n## NIDS and signature-based rule engines\n## Motivation: SOC fatigue and rule maintenance bottlenecks\n## Prior approaches: ML and LLM-driven rule engineering\n# Research Questions\n## Expert trust and deployability of LLM rules\n## Interaction patterns in expert workflows","[{\"question\":\"How do model size and expert perception relate to effectiveness?\",\"answer\":\"Large models (≥70B) consistently generate syntactically valid rules, while small models (≤4B) are largely ineffective; even with a favorable SUS score of 67, practitioners remain skeptical of autonomous generation and prefer LLMs as support tools.\"}]",1784192154,71,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":28},"beyond-the-syntax-do-security-experts-trust-llms-for-nids-rule-engineering","",{"@graph":36,"@context":77},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/beyond-the-syntax-do-security-experts-trust-llms-for-nids-rule-engineering/84033/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"How do model size and expert perception relate to effectiveness?","Question",{"text":75,"@type":76},"Large models (≥70B) consistently generate syntactically valid rules, while small models (≤4B) are largely ineffective; even with a favorable SUS score of 67, practitioners remain skeptical of autonomous generation and prefer LLMs as support tools.","Answer","https://schema.org",{"og:url":51,"og:type":79,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":81,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":84},[85,89,93,97,102,107,112,115,120,123,127],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":46,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":46,"category_name":109,"show_sort_weight":110,"slug":111},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":113,"slug":114},30,"research-report",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},9,"Religion & Spirituality",20,"religion-spirituality",{"id":118,"doc_module":4,"doc_module_name":46,"category_name":121,"show_sort_weight":118,"slug":122},"World Cup","world-cup",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":124,"slug":126},10,"Lifestyle","lifestyle",{"id":128,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":98,"slug":130},19,"General","general"]