Phishing Detector
4-layer cybersecurity pipeline achieving 97% accuracy across English, Hindi, and Hinglish threat content.
What is the Phishing Detector?
An advanced Machine Learning pipeline designed to evaluate URLs, HTML email content, and text to determine the probability of a phishing attack.
Problem
Phishing attacks have evolved to bypass standard heuristic filters. Additionally, many automated detectors fail when analyzing regional or mixed-language threats, such as Hinglish (Hindi + English), which are highly prevalent in the Indian subcontinent.
Results & Evidence
The model was evaluated against a customized, split dataset consisting of verified benign and malicious sources. On the project's evaluated test set, it reports:
* Note: These metrics represent performance on the evaluated dataset and showcase structural ML capabilities rather than enterprise-wide production parity.
Evaluated Benchmark Performance • Multi-Vector Analysis
| Metric | Score | Evaluated Benchmark Scope | Operational Defense Significance |
|---|---|---|---|
| Accuracy | 97% | Comprehensive multi-source test dataset | High generalizability across common threat categories |
| Precision | 96% | Legitimate domain false-positive testing | Minimizes false-alarm disruption for authorized communications |
| Recall | 97% | Obfuscated and look-alike domain evasion | Crucial for zero-day credential harvesting attempts |
| F1-Score | 97% | Multilingual English, Hindi, Hinglish threats | Harmonized balance between sensitivity and specificity |