[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84737-en":3,"doc-seo-84737-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84737,137441390410,"Hazel","https://ap-avatar.wpscdn.com/avatar/2000252f4ab5702993?_k=1776741390130283984",8,"Research & Report","Weakly Guided and Autoregressive Beamformer Parameterization for Generalizable Moving Speaker Extraction in Higher-Order Ambisonics","Linear spatial filters (beamformers) support robust, generalizable and interpretable speech enhancement when parameterization matches ideal assumptions. Modern systems often rely on deep neural parameterization, but performance degrades when multiple moving speakers with unknown directions are present. This work presents a weakly guided, data-driven beamforming pipeline using only an estimate of the target’s initial DoA. A higher-order ambisonics framework decouples neural temporal-spectral processing from linear spatial filtering, enabling array-agnostic generalization with autoregression for consistent results during fast motion and long recordings.","WEAKLY GUIDED AND AUTOREGRESSIVE BEAMFORMER PARAMETERIZATION FOR GENERALIZABLE MOVING SPEAKER EXTRACTION IN HIGHER-ORDER AMBISONICS  \nJakob Kienegger, Tal Peer, Sina Khanagha, Timo Gerkmann  \nSignal Processing (SP) Group, University of Hamburg  \narXiv :2607 .04471v1 [ ee ss .AS] 5 Jul 2026  \nABSTRACT  \nLinear spatial filters (beamformers) enable robust, generalizable and interpretable speech enhancement with performance guarantees under ideal parameterization. Modern beamformers are often parameterized by deep neural networks, whose performance degrades in dynamic scenarios with multiple moving speakers of unknown directions. We propose a data-driven beamforming pipeline, which only requires an estimate of the target’s initial direction. Building on a higher-order ambisonics representation, we show that neural temporal-spectral processing can be decoupled from linear spatial processing, and thereby achieve generalizable and array-agnostic enhancement. By incorporating autoregression into a frame-wise causal framework, we maintain consistent performance throughout fast speaker motion and long recordings. Evaluation on synthetic data demonstrates robust enhancement under challenging conditions with closely spaced and crossing speakers. Real-world recordings in a dynamic office meeting scenario complement these findings and show generalizability across varying ambisonics orders.  \nIndex Terms— Ambisonics, autoregressive, moving speaker, multi-channel speech enhancement, mask-based beamforming.  \n1. INTRODUCTION  \nSpeech enhancement aims to improve the perceptual quality and intelligibility of a recorded speech signal by suppressing noise and reverberation. In a multi-speaker environment, target speaker extraction (TSE) solves the additional challenge of differentiating between the desired target speech and remaining speakers, which are to be treated as noise. When a recording from a microphone array is available, the relative direction of the target to the array, known as direction of arrival (DoA), can be employed to identify the desired speaker. In case of a spherical array, ambisonics [1] provides a directional and array-agnostic representation of the recorded sound field, making it a popular audio format for enhancement [2–5] .  \nAssuming the target’s DoA is known, a spatial filter can be steered toward the desired direction and extract the corresponding speech signal. By jointly processing temporal-spectral and spatial information, recently proposed deep, non-linear spatial filters achieve outstanding enhancement performance [6, 7] . While linear spatial filters, known as beamformers, are proven to be inferior under realistic, non-Gaussian statistical assumptions [8], they provide robustness and interpretability by being a parametric, statistics-based approach. When combined with deep neural networks (DNNs) to  \nThis work was supported by the Deutsche Forschungsgemeinschaft (DFG) under Grant 508337379 . Computational resources were provided by the Regional Computer Center (RRZ) of the University of Hamburg and the Erlangen National High Performance Computing Center (NHR@FAU) under Project f104ac. NHR is funded by the Federal Government and Bavaria. Both facilities received DFG funding under Grants 440719683 and 498394658 .  \nestimate these statistics, such hybrid approaches can achieve strong enhancement performance even with linear spatial filters [9, 10] .  \nIn stationary scenarios, the target’s DoA may be available a priori. However, when speaker motion becomes non-negligible, the continuous directional information necessary to steer a spatial filter—which we refer to as strong guidance—is typically not available. Weakly guided speaker extraction relaxes this constraint and only assumes knowledge of the target’s direction at recording start. To continue using a spatial filter for extraction, an additional tracking algorithm is typically required to infer the target’s movement from the starting direction and automate the s","cbCaimvn11gHotPT","https://ap.wps.com/l/cbCaimvn11gHotPT","pdf",1403167,2,1,5,"English","en",105,"# Abstract\n# Introduction\n## Speech enhancement and target speaker extraction\n## Spatial guidance, weak guidance, and tracking\n## Mask-based beamformer parameterization\n## Causal processing and autoregressive context","[{\"question\":\"What problem does the proposed method address in multi-speaker environments?\",\"answer\":\"It targets moving target speaker extraction when multiple speakers are active and their directions can change over time, making guidance and alignment difficult.\"},{\"question\":\"What information does the method require about the target’s direction?\",\"answer\":\"It only needs an estimate of the target’s initial direction of arrival at the start of the recording, using weak guidance rather than full trajectory knowledge.\"},{\"question\":\"How does the approach maintain performance during fast speaker motion and long recordings?\",\"answer\":\"It incorporates autoregression into a frame-wise causal framework so the method can exploit temporal context consistently across the recording.\"}]",1784197966,13,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"weakly-guided-and-autoregressive-beamformer-parameterization-for-generalizable-moving-speaker-extraction-in-higher-order-ambisonics","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/weakly-guided-and-autoregressive-beamformer-parameterization-for-generalizable-moving-speaker-extraction-in-higher-order-ambisonics/84737/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the proposed method address in multi-speaker environments?","Question",{"text":75,"@type":76},"It targets moving target speaker extraction when multiple speakers are active and their directions can change over time, making guidance and alignment difficult.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What information does the method require about the target’s direction?",{"text":80,"@type":76},"It only needs an estimate of the target’s initial direction of arrival at the start of the recording, using weak guidance rather than full trajectory knowledge.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the approach maintain performance during fast speaker motion and long recordings?",{"text":84,"@type":76},"It incorporates autoregression into a frame-wise causal framework so the method can exploit temporal context consistently across the recording.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,109,114,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":106,"show_sort_weight":107,"slug":108},"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":22,"slug":137},19,"General","general"]