metalalchemistspex commited on
Commit
73b87ae
·
unverified ·
1 Parent(s): a9176fc

Create README.md

Browse files
Files changed (1) hide show
  1. README.md +145 -0
README.md ADDED
@@ -0,0 +1,145 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ title: Smishi / Ne Nasedaj
3
+ emoji: 🎣
4
+ colorFrom: red
5
+ colorTo: orange
6
+ sdk: gradio
7
+ sdk_version: 4.x
8
+ app_file: app.py
9
+ pinned: false
10
+ license: mit
11
+ tags:
12
+ - text-classification
13
+ - phishing-detection
14
+ - serbian
15
+ - bosnian
16
+ - croatian
17
+ - sms-security
18
+ - cybersecurity
19
+ - south-slavic-nlp
20
+ ---
21
+
22
+ # 🎣 Smishi — Ne Nasedaj
23
+
24
+ **SMS Phishing Detector for Serbian, Bosnian, Croatian & Montenegrin**
25
+
26
+ *"Ne Nasedaj" — Don't Fall For It.*
27
+
28
+ ## What is this?
29
+
30
+ Smishi is an open-source SMS phishing (smishing) detector purpose-built for South Slavic languages — Serbian, Bosnian, Croatian, and Montenegrin (SBCM).
31
+
32
+ Paste an SMS message. The model tells you whether it looks like a scam, in either Serbian or English (language toggle in the UI).
33
+
34
+ No app. No login. No data stored. Just paste and check.
35
+
36
+ ## Why does this exist?
37
+
38
+ Smishing is one of the fastest-growing cyber threats in the Western Balkans. Large-scale campaigns impersonating institutions like Pošta Srbije have targeted Serbian citizens for financial fraud and credential theft. Phishing attacks globally have surged sharply since 2022, amplified by AI tools that make localized, convincing fake messages trivially easy to produce at scale.
39
+
40
+ Yet no publicly available smishing detection tool supports SBCM languages. Generic multilingual models struggle with:
41
+
42
+ - **Rich morphological case inflection** — *nagrada / nagradu / nagradi* (prize, accusative, dative) are the same attack keyword in three forms
43
+ - **Cyrillic ↔ Latin script duality** — identical messages can appear in either script
44
+ - **Character substitution evasion** — attackers replace characters (а → @, о → 0) to slip past keyword filters
45
+ - **No existing labeled dataset** — we had to build one from scratch
46
+
47
+ Smishi was built to fill this gap.
48
+
49
+ ## How it works
50
+
51
+ Smishi uses an **ensemble** of two models plus rule-based heuristics:
52
+
53
+ | Component | Detail |
54
+ |---|---|
55
+ | Model A | TF-IDF (character n-grams, 3–5) + Logistic Regression |
56
+ | Model B | Fine-tuned BERTić transformer (`ravi2505/ne-nasedaj-sms-phishing`) |
57
+ | Heuristics | Suspicious/typosquatted domain detection, message-length analysis, threat-vector flags |
58
+ | Training data | 400+ labeled SBCM SMS messages (phishing + legitimate), Cyrillic & Latin |
59
+ | Languages | Serbian (Cyrillic & Latin), Bosnian, Croatian, Montenegrin |
60
+ | Infrastructure | Model A + heuristics run on CPU; Model B uses GPU if available, falls back to CPU |
61
+ | Interface | Bilingual (EN/SR) Gradio app, single-message and batch (CSV) modes |
62
+
63
+ Both model predictions and confidence scores are shown side by side, along with flagged red-flag indicators (suspicious domains, typosquatting, urgency language, etc.).
64
+
65
+ ## Known limitations
66
+
67
+ - **No-URL phishing**: The system currently underperforms on scams that rely on IBAN manipulation, voice phishing (vishing) scripts, or social pressure without links. This is a known gap and active area of improvement.
68
+ - **Dataset size**: ~400 rows is a solid starting point but benefits from more diverse examples, especially Bosnian and Montenegrin regional variants.
69
+ - **Novel attack patterns**: AI-generated smishing may use phrasing outside the training distribution. Contributions welcome.
70
+
71
+ ## Contributing
72
+
73
+ We welcome:
74
+ - New labeled SMS examples (phishing or legitimate) in any SBCM language
75
+ - Edge cases: messages without URLs, IBAN scams, impersonation without links
76
+ - Cyrillic-heavy examples, regional dialect variants
77
+
78
+ Open an issue or pull request on this repository.
79
+
80
+ ## Built by
81
+
82
+ Utaem & Monkeydluffy
83
+ Built during the Build Small Hackathon, June 2026.
84
+
85
+ ## References
86
+
87
+ - UNDP / PwC Serbia, *Tržište rada u oblasti sajber bezbednosti u Srbiji*, Belgrade, June 2026
88
+ - National CERT Serbia, Phishing campaign alert, August 2024
89
+ - M. Drolet, "AI Is Amping Up Phishing, Smishing and Vishing Attacks," Forbes Technology Council, May 2025
90
+ - RATEL, Annual Report 2024
91
+
92
+ ---
93
+
94
+ # 🎣 Smishi — Ne Nasedaj (SR)
95
+
96
+ **Detektor SMS phishinga za srpski, bosanski, hrvatski i crnogorski**
97
+
98
+ *"Ne Nasedaj" — Don't Fall For It.*
99
+
100
+ ## Šta je ovo?
101
+
102
+ Smishi je open-source detektor SMS phishinga (smishinga) napravljen specijalno za južnoslovenske jezike — srpski, bosanski, hrvatski i crnogorski.
103
+
104
+ Zalepi SMS poruku. Model ti kaže da li izgleda kao prevara, na srpskom ili engleskom (toggle u interfejsu).
105
+
106
+ Bez aplikacije. Bez prijave. Bez čuvanja podataka. Samo zalepi i provjeri.
107
+
108
+ ## Zašto postoji?
109
+
110
+ Smishing je jedan od najbrže rastućih sajber pretnji na Zapadnom Balkanu. Kampanje koje imitiraju institucije poput Pošte Srbije ciljaju građane radi finansijske prevare i krađe kredencijala. Phishing napadi globalno su značajno porasli od 2022. godine, pojačani AI alatima koji lokalizovane, uverljive lažne poruke čine trivijalnim za masovnu produkciju.
111
+
112
+ Ipak, ne postoji javno dostupan alat za detekciju smishinga koji podržava naše jezike. Generički višejezični modeli imaju problema sa:
113
+
114
+ - **Bogatom morfološkom fleksijom** — *nagrada / nagradu / nagradi* su ista ključna reč napada u tri oblika
115
+ - **Ćirilica ↔ latinica dualnost** — ista poruka može biti napisana u oba pisma
116
+ - **Supstitucijom karaktera** — napadači zamenjuju slova (а → @, о → 0) da prevare filtere
117
+ - **Nepostojanjem označenog dataseta** — morali smo da ga izgradimo od nule
118
+
119
+ ## Kako radi
120
+
121
+ Smishi koristi **ensemble** dva modela plus heuristike zasnovane na pravilima:
122
+
123
+ | Komponenta | Detalj |
124
+ |---|---|
125
+ | Model A | TF-IDF (karakterski n-grami, 3–5) + Logistička regresija |
126
+ | Model B | Fine-tuned BERTić transformer (`ravi2505/ne-nasedaj-sms-phishing`) |
127
+ | Heuristike | Detekcija sumnjivih/typosquat domena, analiza dužine poruke, threat-vector indikatori |
128
+ | Podaci | 400+ označenih SBCM SMS poruka (phishing + legitimne), ćirilica i latinica |
129
+ | Jezici | Srpski (ćirilica i latinica), bosanski, hrvatski, crnogorski |
130
+ | Infrastruktura | Model A + heuristike na CPU-u; Model B koristi GPU ako je dostupan, inače CPU |
131
+ | Interfejs | Dvojezični (EN/SR) Gradio app, pojedinačni i batch (CSV) režim |
132
+
133
+ ## Poznata ograničenja
134
+
135
+ - **Phishing bez URL-a**: Sistem trenutno slabije prepoznaje prevare koje se oslanjaju na IBAN manipulaciju, vishing skripte ili socijalni pritisak bez linkova.
136
+ - **Veličina dataseta**: ~400 redova je dobra osnova, ali model profitira od više primera, posebno iz bosanskih i crnogorskih regionalnih varijanti.
137
+
138
+ ## Doprinos
139
+
140
+ Dobrodošli su novi označeni SMS primeri, granični slučajevi (bez URL-a, IBAN prevare, imitacija bez linkova), ćirilični i regionalni dijalekatski primeri. Otvori issue ili pull request.
141
+
142
+ ## Napravili
143
+
144
+ Utaem & Monkeydluffy
145
+ Napravljeno tokom Build Small Hackathona, juni 2026.