[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-122821-en":3,"doc-seo-122821-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},122821,3848291630094,"Emma Wilson","https://eur-avatar.wpscdn.com/davatar_085a072bc5b1113ac321206ff7593b45",6,"Technology","Personalized Mixed Reality Audio Using Audio Classification Using Machine Learning","Traditional audio playback through over-the-ear or in-ear headphones can be limited in noisy environments, making it hard to focus on the most important sound. This disclosure presents a mixed-reality audio approach that uses machine learning to modify audio according to user preferences, such as isolating a single speaker in a group, reducing competing chatter, and enhancing muffled speech. With user permission, notifications can also be personalized and rendered in a voice and tone aligned to the message sender for a more natural, engaging listening experience.","Technical Disclosure Commons  \nDefensive Publications Series  \nDecember 2023  \nPersonalized Mixed Reality Audio Using Audio Classification Using Machine Learning  \nQuinn Thuy Tran Joseph Johnson Jr.  \nFollow this and additional works at: [https://www.tdcommons.org/dpubs_series](https://www.tdcommons.org/dpubs_series)  \nRecommended Citation  \nTran, Quinn Thuy and Johnson Jr., Joseph, \"Personalized Mixed Reality Audio Using Audio Classification Using Machine Learning\", Technical Disclosure Commons,(December 15, 2023)  \n[https://www.tdcommons.org/dpubs_series/6496](https://www.tdcommons.org/dpubs_series/6496)  \nThis work is licensed under a Creative Commons Attribution 4.0 License.  \nThis Article is brought to you for free and open access by Technical Disclosure Commons. It has been accepted for inclusion in Defensive Publications Series by an authorized administrator of Technical Disclosure Commons.  \nPersonalized Mixed Reality Audio Using Audio Classification Using Machine Learning  \nABSTRACT  \nTraditional audio devices such as over-the-ear or in-ear headphones are limited in their  \nability to provide a personalized and engaging listening experience. When using such audio  \ndevices in a noisy environment, it can be difficult for a user to focus on the audio that is most  \nimportant. This disclosure describes the use of mixed reality audio to enhance the listening  \nexperience via headphones, earbuds, or other audio devices during normal use. Machine learning techniques are used to modify the audio per user preferences, e.g., to focus the audio on a single  \nperson talking while the user is in a group, to turn down noisy/competing audio such as other  \npeople talking in a busy/noisy group setting like a crowd or party, to enhance muffled words  \nwith clean versions, etc. The modified audio is played back via headphones, earbuds, or other  \ndevice to provide a personalized listening experience.  \nKEYWORDS  \n● Generative audio  \n● Personalized audio  \n● Immersive audio  \n● Audio enhancement  \n● Mixed reality  \n● Voice model  \n● Hearability  \n● Directional audio  \n● Muffled audio  \nPublished by Technical Disclosure Commons, 2023 2  \nBACKGROUND  \nTraditional audio devices such as over-the-ear or in-ear headphones are limited in their  \nability to provide a personalized and engaging listening experience. For example, when using  \nsuch audio devices in a noisy environment, it can be difficult for a user to focus on the audio that is most important. Additionally, messages or notifications received on a personal device such asa smartphone are read aloud (via speech-to-text) in a generic voice, and delivery of such audio is  \nnot personalized for the user.  \nDESCRIPTION  \nThis disclosure describes the use of mixed reality audio to enhance the listening experience via headphones, earbuds, or other audio devices during normal use. Machine learning techniques are used to modify the audio per user preferences, e.g., to focus the audio on a single person talking while the user is in a group, to turn down noisy/competing audio such as other  \npeople talking in a busy/noisy group setting like a crowd or party, to enhance muffled words  \nwith clean versions, etc. The modified audio is played back via headphones, earbuds, or other  \ndevice to provide a personalized listening experience.  \nMachine learning techniques are used to identify audio by object or category, e.g., sounds of birds, sounds of water or rain, etc. and isolate such audio. Image-based neural networks are available that can perform this task on images. These networks can also be trained with images of Mel Frequency Cepstrum Coefficient (MFCC) feature vectors for audio including speech and categorized in the same manner. Alternatively, modern GPT-based multimodal foundational  \nmodels that are trained on image, text, and audio data can be utilized for this purpose.  \nAfter such identification, the audio is modified per user preference, e.g., made louder or softer. For example, su","cbCaiuuU6QGq9W8E","https://ap.wps.com/l/cbCaiuuU6QGq9W8E","pdf",217024,1,9,"English","en",105,"# Abstract\n# Background\n# Description\n## Audio identification and isolation\n## Preference-based audio modification\n## User-permission personalization for notifications\n## Example mixed reality smart sound selection and enhancement process","[{\"question\":\"How does the disclosure improve listening in noisy environments?\",\"answer\":\"It uses mixed reality audio enhanced by machine learning to modify sounds based on user preferences, such as focusing on a specific person and reducing competing audio like other people talking in a crowd.\"},{\"question\":\"What machine learning methods are used for identifying audio?\",\"answer\":\"The disclosure describes using audio classification via object/category isolation, including image-based neural networks trained on MFCC feature vectors for speech and categorized audio, or GPT-based multimodal foundational models trained on image, text, and audio data.\"},{\"question\":\"Can the approach personalize device notifications as well as ambient audio?\",\"answer\":\"Yes. With user permission, notifications such as text messages can be read aloud using the sender’s voice and tone, incorporating the user’s style and the message content to sound more natural.\"}]","Personalized Mixed Reality Audio Using Audio Classification Using Machine Learning | PDF",1785813083,23,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"personalized-mixed-reality-audio-using-audio-classification-using-machine-learning","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/technology/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/personalized-mixed-reality-audio-using-audio-classification-using-machine-learning/122821/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"How does the disclosure improve listening in noisy environments?","Question",{"text":75,"@type":76},"It uses mixed reality audio enhanced by machine learning to modify sounds based on user preferences, such as focusing on a specific person and reducing competing audio like other people talking in a crowd.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What machine learning methods are used for identifying audio?",{"text":80,"@type":76},"The disclosure describes using audio classification via object/category isolation, including image-based neural networks trained on MFCC feature vectors for speech and categorized audio, or GPT-based multimodal foundational models trained on image, text, and audio data.",{"name":82,"@type":73,"acceptedAnswer":83},"Can the approach personalize device notifications as well as ambient audio?",{"text":84,"@type":76},"Yes. With user permission, notifications such as text messages can be read aloud using the sender’s voice and tone, incorporating the user’s style and the message content to sound more natural.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,113,118,123,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":111,"slug":112},50,"technology",{"id":114,"doc_module":4,"doc_module_name":46,"category_name":115,"show_sort_weight":116,"slug":117},7,"Healthcare",40,"healthcare",{"id":119,"doc_module":4,"doc_module_name":46,"category_name":120,"show_sort_weight":121,"slug":122},8,"Research & Report",30,"research-report",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]