Abstract
Tor is widely used to protect users’ anonymity online by mitigating IP-based tracking and blocking browser-level fingerprinting techniques. However, while the Tor browser effectively prevents passive data collection, it cannot stop users from voluntarily disclosing identifying information in web forms provided by onion services. This user-mediated information leakage represents a significant and underexplored threat vector to anonymity. In this work, we propose a fully client-side approach based on Natural Language Processing (NLP) to extend the Tor Browser with the ability to detect potentially deanonymizing information requests in real time. Our method analyzes webpage content—particularly form labels—to identify and flag attempts to solicit personally identifying data. This approach complements existing Tor browser protections by extending them to the semantic level of human-browser interaction. We validate our system using a dataset of real onion service websites, from which we automatically extract form-based user prompts. Interestingly, our analysis shows that 44% of the examined onion services requiring user information include requests for personal data that may compromise user anonymity. Our method achieves over 90% on all standard classification metrics (precision, accuracy, recall, and F1-score) in detecting whether a website requests information that could potentially deanonymize the user, while introducing a negligible processing overhead that does not impact the user’s browsing experience.
Introduction
Anonymity in online communication remains a fundamental requirement for individuals operating under surveillance or censorship [2, 8, 13]. Tor [5] is one of the most widely adopted protocols to safeguard user anonymity by concealing IP addresses. Notably, the Tor Browser is designed to thwart common fingerprinting [7] and tracking techniques employed by websites, such as accessing the local timezone or language settings via JavaScript APIs. For instance, any attempt by a script to retrieve the current time and timezone using a command like new Date() will be sandboxed: the browser will override the returned values to reflect a fixed GMT timezone, preventing sites from inferring the user’s geographical location.
Despite these robust browser-level protections, one critical vulnerability remains unresolved: the user. While the browser can block passive fingerprinting and data leakage attempts, it cannot prevent users from actively — and often unknowingly — divulging personal information. For example, many onion services may request users to fill out forms containing sensitive or identifying information such as real names, email addresses, physical locations, or subtly revealing details like preferred language or country of residence. These interactions occur within dynamic and hypertextual interfaces, where human-computer interaction becomes a potential attack vector, bypassing browser-level anonymity guarantees.
This paper introduces a novel mitigation strategy to address this overlooked vulnerability. We propose an NLP-based solution that can operate entirely on the client side (for example, by extending the Tor browser) and is capable of detecting potentially deanonymizing fields on onion service websites. Our solution semantically analyzes the hypertextual structure of web pages in real-time, flagging requests for sensitive information that could compromise the user’s anonymity if voluntarily disclosed. This approach complements existing browser protections by extending them to the semantic level of human-browser interaction.
Our approach is designed to be lightweight and operate at runtime during web navigation, ensuring minimal to no impact on the user experience. Moreover, the proposed solution does not rely on external services or APIs that could collect user browsing data or require subscriptions, thereby preserving user anonymity.
To the best of our knowledge, no current solution integrates semantic content analysis within the Tor Browser to guard against voluntary data leakage by users. While previous work has focused on preventing passive information leaks [1, 6] (e.g., via browser APIs or timing attacks), our approach addresses the deanonymization problem due to human-browser interaction.
In our evaluation, we analyzed a large sample of onion services and found that approximately 44% of onion sites requiring form filling, include form fields that request personal data potentially compromising user anonymity. Our NLP-based detection method demonstrates high effectiveness, achieving over 90% across standard classification metrics — including precision, recall, accuracy, and F1-score — while incurring minimal runtime overhead.
The remainder of this paper is organized as follows. Section 2 provides an overview of related work Section 3 describes the design of our proposed solution, detailing the pipeline for identifying anonymity-compromising form fields on Onion websites. In Section 4, we present the experimental setup, dataset collection process, and the results obtained by applying our methodology, including performance metrics and efficiency analysis. Finally, Section 5 concludes the paper and outlines directions for future works.
Related Work
The Tor network is widely adopted for online anonymity, and considerable research has been dedicated to reinforcing its privacy guarantees against both passive and active deanonymization attacks. Prior efforts have predominantly targeted side-channel attacks [4], network-level traffic correlation attacks [3, 10] and browser-level timing attacks [1].
Tools like Tor Browser have addressed many de-anonymization attempts through sandboxing and uniform behavior enforcement [14]. However, less attention has been given to active information leakage, where users voluntarily disclose personal data via web forms. Notably, prior work on deanonymization risks has focused on adversarial JavaScript execution [1], or misuse of browser APIs [6], but has largely overlooked semantic-level risks introduced by onion service interfaces themselves.
Recent advances in Natural Language Processing (NLP) and the emergence of lightweight embedding models, such as Sentence-BERT [11], have enabled practical semantic analysis directly in the browser. While some researchers employ NLP to analyze web browsing [9, 12], none, to our knowledge, have been designed specifically to flag deanonymizing prompts in the context of privacy networks like Tor.
Our work uniquely addresses this gap by proposing an NLP-based, client-side approach that detects anonymity-compromising form fields using semantic similarity, without relying on external infrastructure.
Methodology
Our proposed solution enhances user anonymity during navigation on Tor Browser, particularly addressing the risk of partial or full de-anonymization when users interact with form fields on Onion websites. The primary objective of our approach is to identify form fields that potentially compromise anonymity, efficiently and effectively, without introducing significant latency to the user’s browsing experience.
The methodology consists of the following pipeline.
Step 1: Keyword Generation. Initially, a comprehensive set of keywords is generated using the GPT-4 Large Language Model (LLM). The model is prompted to provide relevant keywords related to directly identifying information (e.g., email, phone number), pseudonymous identifiers (e.g., usernames, passwords), location-based data (e.g., country, language), and other metadata capable of compromising anonymity.
Figure 1 reports the LLM prompt employed for generating these keywords.
This generated keyword list is computed once and serves as the reference for subsequent similarity checks.
This approach was chosen to leverage the LLM’s broad contextual knowledge and linguistic flexibility, enabling the generation of a diverse and semantically rich keyword set beyond what manual curation or rule-based methods typically offer.
Step 2: Keyword Embeddings Generation. Subsequently, embeddings for the generated keyword list are computed using the SentenceTransformer model all-MiniLM-L6-v2. This embedding representation facilitates efficient semantic comparison between keywords and form fields during user navigation.
Step 3: Form Field Extraction. When a user visits an Onion website, the HTML content of the page is automatically extracted for further analysis. From the extracted HTML, all visible and interactive form fields are identified and parsed using the Beautiful Soup library. Specifically, hidden input fields and fields with CSS attribute style="displaynone" are excluded from analysis to avoid false positives.
Step 4: Similarity Computation and Risk Identification. For each extracted form field, embeddings are calculated using the same SentenceTransformer model. Cosine similarity scores between each form field embedding and the precomputed keyword embeddings are then computed. If a similarity score exceeds a predefined threshold, the corresponding form field is flagged as potentially anonymity-compromising, consequently marking the entire site as posing a possible anonymity risk. As an heuristic approach, if a form field is a substring of at least 4 characters of a keyword in the precomputed list, then it is classified as potentially anonymity-compromising.
A schematic flowchart representation of our pipeline is depicted in Figure 2. This visual representation clarifies the sequential flow and integration between each step.
Validation and Results
In this section, we describe the dataset collection and the results of our experiments.
Dataset Collection and Ground Truth Annotation
On May 17, 2025, we collected a dataset of Onion sites by scraping results from https://ahmia.fi/search/?q=, resulting in 577 unique Onion addresses. Ahmia is an open-source search engine for Tor hidden services (i.e.,.onion sites), developed with support from the Tor Project. It indexes publicly accessible Onion addresses while filtering out illegal content, making it a reliable and ethical resource for discovering active services on the dark web.
We attempted to connect to each of these 577 Onion addresses using an automated scraper and successfully retrieved HTML content from 443 active Onion sites.
From these 443 active sites, we extracted all form fields by parsing the HTML and isolating user input forms, excluding hidden and non-visible elements. A total of 117 sites were found to contain forms, yielding a combined dataset of 442 individual form fields.
Each of the 442 form fields was manually labeled based on whether it could potentially compromise user anonymity. This manual annotation was conducted by visiting all 117 sites and considering contextual elements such as prefilled inputs, placeholder texts, and surrounding content. This process resulted in a labeled dataset to serve as the ground truth for evaluating our pipeline. We found that 44% of 117 onion sites containing forms include requests for personal data that may compromise user anonymity.
Results
We evaluated our pipeline on the annotated dataset using standard classification metrics: Accuracy, Precision, Recall, and F1-Score.
Figures 3 and 4 show the metric performance curves across varying cosine similarity thresholds. Figure 3 corresponds to site-level classification (i.e., identifying sites as Privacy-Leaking vs. Safe), while Figure 4 refers to field-level classification (i.e., identifying form fields as anonymity-compromising or not).
Both curves indicate that a threshold of 0.74 provides a balanced trade-off among the four metrics. Lower thresholds increase recall but reduce precision, while higher thresholds achieve the opposite. Notably, false negatives may occur due to obfuscated or non-standard field labels—e.g., a username field labeled simply as "u". Such cases are rare, as reflected in the overall performance.
Table 1: Performance metrics at optimal threshold (0.74)
Classifier | Accuracy | Precision | Recall | F1-Score |
|---|---|---|---|---|
Site-level | 0.9316 | 0.9400 | 0.9038 | 0.9216 |
Field-level | 0.9389 | 0.9018 | 0.8632 | 0.8821 |
Figures 5 and 6 display the ROC curves for the site and field classifiers respectively. The Area Under the Curve (AUC) values of 0.93 and 0.91 further support the robustness of our classification approach.
Performance Evaluation
We also assessed the computational efficiency of our pipeline. On a standard machine equipped with an Intel(R) Core(TM) i7-10510U CPU @ 1.80GHz and 16GB RAM:
Embedding generation for a list of 180 keywords takes approximately 0.2–0.3 seconds (performed once).
Full site classification (including HTML parsing, embedding, and similarity checks) takes 0.01–0.03 seconds.
These timings confirm that our pipeline introduces negligible overhead to the browsing experience, making it suitable for real-time applications in privacy-enhancing technologies.
Conclusion
While the Tor Browser implements robust mechanisms to prevent passive deanonymization via browser fingerprinting and network-level tracking, it remains vulnerable to a critical and under-addressed threat vector: user-driven information disclosure. In this paper, we introduce an NLP-based solution designed to operate entirely on the client side within the Tor Browser. Our contribution complements existing browser-level protections by shifting the focus from system-level threats to user-facing content and interactions.
Through evaluation on a dataset of real-world onion sites, we demonstrated the effectiveness of our approach in detecting de-anonymizing input requests, both in terms of classification accuracy and time performance. Our results show that even within the anonymous ecosystem of Tor, many services solicit data that — if provided by the user — could significantly weaken or entirely compromise anonymity. As future work, we plan to develop a fully integrated browser extension tailored for the Tor Browser. This integration should facilitate real-time notifications and warnings in a way that aligns with the Tor Browser’s existing design principles and usability constraints. Another direction involves extending the system’s multilingual capabilities. Currently, our NLP pipeline only works with English-language content. To improve coverage, our integration will incorporate multilingual language models and translation-aware processing.
Acknowledgments
This work is partially supported by project SERICS (PE00000014) under the MUR National Recovery and Resilience Plan funded by the European Union - NextGenerationEU.
References
[1] Timothy G Abbott, Katherine J Lai, Michael R Lieberman, and Eric C Price. 2007. Browser-based attacks on Tor. In International Workshop on Privacy Enhancing Technologies. Springer, 184–199.
[2] Francesco Buccafurri, Vincenzo de Angelis, and Sara Lazzaro. 2023. MQTT-A: A Broker-Bridging P2P Architecture to Achieve Anonymity in MQTT. IEEE Internet of Things Journal 10, 17 (2023), 15443–15463. 10.1109/JIOT.2023.3264019
[3] Giovanni Cherubin, Rob Jansen, and Carmela Troncoso. 2022. Online website fingerprinting: Evaluating website fingerprinting attacks on tor in the real world. In 31st USENIX Security Symposium (USENIX Security 22). 753–770.
[4] Mila Dalla Preda, Claudia Greco, Michele Ianni, Francesco Lupia, and Andrea Pugliese. 2025. Light sensor based covert channels on mobile devices. Information Sciences 690 (2025), 121581. 10.1016/j.ins.2024.121581
[5] Roger Dingledine, Nick Mathewson, Paul F Syverson, et al. 2004. Tor: The second-generation onion router.. In USENIX security symposium , Vol. 4. 303–320.
[6] Pierre Laperdrix, Nataliia Bielova, Benoit Baudry, and Gildas Avoine. 2020. Browser Fingerprinting: A Survey. ACM Trans. Web 14, 2, Article 8 (April 2020), 33 pages. 10.1145/3386040
[7] Mohamad Amar Irsyad Mohd Aminuddin, Zarul Fitri Zaaba, Azman Samsudin, Faiz Zaki, and Nor Badrul Anuar. 2023. The rise of website fingerprinting on Tor: Analysis on techniques and assumptions. Journal of Network and Computer Applications 212 (2023), 103582. 10.1016/j.jnca.2023.103582
[8] Mainack Mondal, Denzil Correa, and Fabrício Benevenuto. 2020. Anonymity Effects: A Large-Scale Dataset from an Anonymous Social Media Platform. In Proceedings of the 31st ACM Conference on Hypertext and Social Media (Virtual Event, USA) (HT ’20). Association for Computing Machinery, New York, NY, USA, 69–74. 10.1145/3372923.3404792
[9] Daniel Perdices, Javier Ramos, Jose L Garcia-Dorado, Ivan Gonzalez, and Jorge E López de Vergara. 2021. Natural language processing for web browsing analytics: Challenges, lessons learned, and opportunities. Computer Networks 198 (2021), 108357.
[10] Florian Platzer, Marcel Schäfer, and Martin Steinebach. 2020. Critical traffic analysis on the tor network. In Proceedings of the 15th International Conference on Availability, Reliability and Security (Virtual Event, Ireland) (ARES ’20). Association for Computing Machinery, New York, NY, USA, Article 77, 10 pages. 10.1145/3407023.3409180
[11] Nils Reimers and Iryna Gurevych. 2019. Sentence-bert: Sentence embeddings using siamese bert-networks. arXiv preprint arXiv:https://arXiv.org/abs/1908.10084 (2019).
[12] Ozgur Koray Sahingoz, Ebubekir Buber, Onder Demir, and Banu Diri. 2019. Machine learning based phishing detection from URLs. Expert Systems with Applications 117 (2019), 345–357.
[13] Fatemeh Shirazi, Milivoj Simeonovski, Muhammad Rizwan Asghar, Michael Backes, and Claudia Diaz. 2018. A Survey on Routing in Anonymous Communication Protocols. ACM Comput. Surv. 51, 3, Article 51 (June 2018), 39 pages. 10.1145/3182658
[14] Peter Story, Daniel Smullen, Rex Chen, Yaxing Yao, Alessandro Acquisti, Lorrie Faith Cranor, Norman Sadeh, and Florian Schaub. 2022. Increasing adoption of tor browser using informational and planning nudges. Proceedings on Privacy Enhancing Technologies (2022).
Do you like what you are reading? Subscribe to receive updates.
Unsubscribe anytime