| BlackWeb is a project that collects and unifies public blocklists of domains (porn, downloads, drugs, malware, spyware, trackers, bots, social networks, warez, weapons, etc.) to make them compatible with Squid-Cache. | BlackWeb es un proyecto que recopila y unifica listas públicas de bloqueo de dominios (porno, descargas, drogas, malware, spyware, trackers, bots, redes sociales, warez, armas, etc.) para hacerlas compatibles con Squid-Cache. |
squid(orsquid-openssl) (for Squid Rule)
apt install -y squidbwupdate/
├── bwupdate.sh # Main update script: downloads blocklists, builds blackweb.txt, reloads Squid
├── lst/ # ACL/support lists used by bwupdate.sh
│ ├── ai.txt # separate Squid acl: AI-related domains, not touched by bwupdate.sh
│ ├── allowdomains.txt # separate Squid acl: excludes essential domains, not touched by bwupdate.sh
│ ├── allowtlds.txt # required for the debug step: allowed-TLD pattern
│ ├── blockdomains.txt # separate Squid acl: blocks domains outside blackweb.txt, not touched by bwupdate.sh
│ ├── blocktlds.txt # separate Squid acl: blocks whole TLDs, not touched by bwupdate.sh
│ ├── debugbl.txt # merged into the candidate blocklist before debugging
│ ├── debugwl.txt # excludes false positives during debugging (google, hotmail, yahoo, etc.)
│ ├── invalid.txt # not read by any script, maintainer reference
│ ├── pendingtld.txt # not read by any script, maintainer reference
│ ├── streaming.txt # separate Squid acl: streaming domains, not touched by bwupdate.sh
│ ├── tldsappx.txt # TLD source read by dofi/domfilter.py
│ └── tldsbk.txt # not read by any script, maintainer reference
└── tools/
├── checksources.sh # Downloads all source lists and searches them for a domain
└── debugerror.py # Removes domains flagged by Squid's own cache.log errors whose parent is already blocked
| ACL | Blocked Domains | File Size |
|---|---|---|
| blackweb.txt | 7963236 | 194,5 MB |
git clone --depth=1 https://github.com/maravento/blackweb.git
blackweb.txt is already updated and optimized for Squid-Cache. Download it and unzip it in the path of your preference and activate the BlackWeb Rule for Squid-Cache.
|
blackweb.txt ya viene actualizada y optimizada para Squid-Cache. Descárguela y descomprímala en la ruta de su preferencia y active la BlackWeb Rule for Squid-Cache.
|
wget -q -c -N https://raw.githubusercontent.com/maravento/blackweb/master/blackweb.tar.gz && cat blackweb.tar.gz* | tar xzf -#!/bin/bash
# Variables
url="https://raw.githubusercontent.com/maravento/blackweb/master/blackweb.tar.gz"
wgetd="wget -q -c --timestamping --no-check-certificate --retry-connrefused --timeout=10 --tries=4 --show-progress"
# TMP folder
output_dir="bwtmp"
mkdir -p "$output_dir"
# Download
if $wgetd "$url"; then
echo "File downloaded: $(basename $url)"
else
echo "Main file not found. Searching for multiparts..."
# Multiparts from a to z
all_parts_downloaded=true
for part in {a..z}{a..z}; do
part_url="${url%.*}.$part"
if $wgetd "$part_url"; then
echo "Part downloaded: $(basename $part_url)"
else
echo "Part not found: $part"
all_parts_downloaded=false
break
fi
done
if $all_parts_downloaded; then
# Rebuild the original file in the current directory
cat blackweb.tar.gz.* > blackweb.tar.gz
echo "Multipart file rebuilt"
else
echo "Multipart process cannot be completed"
exit 1
fi
fi
# Unzip the file to the output folder
tar -xzf blackweb.tar.gz -C "$output_dir"
echo "Done"wget -q -c -N https://raw.githubusercontent.com/maravento/blackweb/master/blackweb.tar.gz && cat blackweb.tar.gz* | tar xzf -
wget -q -c -N https://raw.githubusercontent.com/maravento/blackweb/master/blackweb.txt.sha256
LOCAL=$(sha256sum blackweb.txt | awk '{print $1}'); REMOTE=$(awk '{print $1}' blackweb.txt.sha256); echo "$LOCAL" && echo "$REMOTE" && [ "$LOCAL" = "$REMOTE" ] && echo OK || echo FAILBlackWeb Rule for Squid-Cache
| Edit: | Edite: |
/etc/squid/squid.conf| And add the following lines: | Y agregue las siguientes líneas: |
# INSERT YOUR OWN RULE(S) HERE TO ALLOW ACCESS FROM YOUR CLIENTS
# Block Rule for Blackweb
acl blackweb dstdomain "/path_to/blackweb.txt"
http_access deny blackweb| BlackWeb contains millions of domains, therefore it is recommended: | BlackWeb contiene millones de dominios, por tanto se recomienda: |
Use allowdomains.txt to exclude essential domains or subdomains, such as .accounts.google.com, .yahoo.com or .github.com. According to Squid's documentation, Google may use the subdomains accounts.google.com and accounts.youtube.com for authentication within its ecosystem. Blocking them could disrupt access to services such as Gmail, Drive and Docs.
|
Use allowdomains.txt para excluir dominios o subdominios esenciales, como .accounts.google.com, .yahoo.com o .github.com. Según la documentación de Squid, Google puede utilizar los subdominios accounts.google.com y accounts.youtube.com para la autenticación dentro de su ecosistema. Bloquearlos podría interrumpir el acceso a servicios como Gmail, Drive y Docs.
|
acl allowdomains dstdomain "/path_to/allowdomains.txt"
http_access allow allowdomains
Use blockdomains.txt to block any other domain not included in blackweb.txt
|
Utilice blockdomains.txt para bloquear cualquier otro dominio no incluido en blackweb.txt
|
acl blockdomains dstdomain "/path_to/blockdomains.txt"
http_access deny blockdomains
Use blocktlds.txt to block gTLD, sTLD, ccTLD, etc.
|
Use blocktlds.txt para bloquear gTLD, sTLD, ccTLD, etc.
|
acl blocktlds dstdomain "/path_to/blocktlds.txt"
http_access deny blocktldsInput:
.bardomain.xxx
.subdomain.bardomain.xxx
.bardomain.ru
.bardomain.adult
.foodomain.com
.foodomain.pornOutput:
.foodomain.com| Use this rule to block Punycode - RFC3492, IDN | Non-ASCII (TLDs or Domains), to prevent an IDN homograph attack. For more information visit welivesecurity: Homograph attacks. | Usar esta regla para bloquear Punycode - RFC3492, IDN | Non-ASCII (TLDs o Dominios), para prevenir un ataque homógrafo IDN. Para mayor información visite welivesecurity: Ataques homográficos. |
acl punycode dstdom_regex -i \.xn--.*
http_access deny punycodeInput:
.bücher.com
.mañana.com
.google.com
.auth.wikimedia.org
.xn--fiqz9s
.xn--p1aiASCII Output:
.google.com
.auth.wikimedia.org| Use this rule to block patterns (Optional. Can generate false positives). | Use esta regla para bloquear patrones (Opcional. Puede generar falsos positivos). |
# Example: Download ACL:
sudo wget -P /etc/acl/squid https://raw.githubusercontent.com/maravento/vault/refs/heads/master/gateproxy/acl/squid/blockpatterns.txt
# Squid Rule to Block Patterns (change path):
acl blockwords url_regex -i "/etc/acl/squid/blockpatterns.txt"
http_access deny blockpatternsInput:
.bittorrent.com
https://www.google.com/search?q=torrent
https://www.google.com/search?q=mydomain
https://www.google.com/search?q=porn
.mydomain.comOutput:
https://www.google.com/search?q=mydomain
.mydomain.com
Use streaming.txt to block (or allow) streaming domains not included in blackweb.txt (for example: .youtube.com, .googlevideo.com, .ytimg.com, etc.).
|
Utilice streaming.txt para bloquear (o permitir) dominios de streaming no incluidos en blackweb.txt (por ejemplo: .youtube.com, .googlevideo.com, .ytimg.com, etc.).
|
acl streaming dstdomain "/path_to/streaming.txt"
http_access deny streaming
Note: This list may contain overlapping domains. It is important to manually clean it according to the proposed objective. Example:
|
Nota: Esta lista puede contener dominios superpuestos. Es importante depurarla manualmente según el objetivo propuesto. Ejemplo:
|
# Block Facebook
.fbcdn.net
.facebook.com
# Block some Facebook streaming content
.z-p3-video.flpb1-1.fna.fbcdn.net
Use ai.txt to block (or allow) domains related to artificial intelligence AI.
|
Utilice ai.txt para bloquear (o permitir) dominios relacionados con inteligencia artificial IA.
|
acl ai dstdomain "/path_to/ai.txt"
http_access deny ai# INSERT YOUR OWN RULE(S) HERE TO ALLOW ACCESS FROM YOUR CLIENTS
# Allow Rule for Domains
# https://raw.githubusercontent.com/maravento/blackweb/master/bwupdate/lst/allowdomains.txt
acl allowdomains dstdomain "/path_to/allowdomains.txt"
http_access allow allowdomains
# Block Rule for Punycode
acl punycode dstdom_regex -i \.xn--.*
http_access deny punycode
# Block Rule for gTLD, sTLD, ccTLD
# https://raw.githubusercontent.com/maravento/blackweb/master/bwupdate/lst/blocktlds.txt
acl blocktlds dstdomain "/path_to/blocktlds.txt"
http_access deny blocktlds
# Block Rule for Domains
# https://raw.githubusercontent.com/maravento/blackweb/master/bwupdate/lst/blockdomains.txt
acl blockdomains dstdomain "/path_to/blockdomains.txt"
http_access deny blockdomains
# Block Rule for Patterns (Optional)
# https://raw.githubusercontent.com/maravento/vault/refs/heads/master/gateproxy/acl/squid/blockpatterns.txt
acl blockwords url_regex -i "/path_to/blockpatterns.txt"
http_access deny blockpatterns
# Block Rule for Blackweb
# https://raw.githubusercontent.com/maravento/blackweb/master/blackweb.tar.gz
acl blackweb dstdomain "/path_to/blackweb.txt"
http_access deny blackweb| This section is only to explain how the update and optimization process works. It is not necessary for the user to run it. This process can take time and consume a lot of hardware and bandwidth resources, therefore it is recommended to use test equipment. | Esta sección es únicamente para explicar cómo funciona el proceso de actualización y optimización. No es necesario que el usuario la ejecute. Este proceso puede tardar y consumir muchos recursos de hardware y ancho de banda, por tanto se recomienda usar equipos de pruebas. |
Note: bwupdate.sh tested on Ubuntu 24.04/26.04 LTS. Use on other versions or distributions is at your own risk.
|
Nota: bwupdate.sh ha sido probado en Ubuntu 24.04/26.04 LTS. Su uso en otras versiones o distribuciones queda bajo su propio riesgo.
|
- Python 3.x, Bash 5.x
wget,curl,tar,gzip,idn2,squid(orsquid-openssl),python3,bind9-host,findutils,grep,sed,coreutils,util-linux,file,libc-bin,sudo- Python module:
requests(required bydomfilter.py)
apt install -y wget curl tar gzip idn2 squid python3 bind9-host findutils grep sed coreutils python3-requests util-linux file libc-bin sudo
The update process of blackweb.txt consists of several steps and is executed in sequence by the script bwupdate.sh. The script will request privileges when required.
|
El proceso de actualización de blackweb.txt consta de varios pasos y es ejecutado en secuencia por el script bwupdate.sh. El script solicitará privilegios cuando lo requiera.
|
git clone --depth=1 https://github.com/maravento/blackweb.git && cd blackweb/bwupdate && ./bwupdate.sh
Make sure your Squid is installed correctly. If you have any problems, run the following script: (sudo ./squid_install.sh):
|
Asegúrese de que Squid esté instalado correctamente. Si tiene algún problema, ejecute el siguiente script: (sudo ./squid_install.sh):
|
#!/bin/bash
# kill old version
while pgrep squid > /dev/null; do
echo "Waiting for Squid to stop..."
killall -s SIGTERM squid &>/dev/null
sleep 5
done
# squid remove (if exist)
apt purge -y squid* &>/dev/null
rm -rf /var/spool/squid* /var/log/squid* /etc/squid* /dev/shm/* &>/dev/null
# squid install (you can use 'squid-openssl' or 'squid')
apt install -y squid-openssl squid-langpack squid-common squidclient squid-purge
# create log
if [ ! -d /var/log/squid ]; then
mkdir -p /var/log/squid
fi &>/dev/null
if [[ ! -f /var/log/squid/{access,cache,store,deny}.log ]]; then
touch /var/log/squid/{access,cache,store,deny}.log
fi &>/dev/null
# permissions
chown -R proxy:proxy /var/log/squid
# enable service
systemctl enable squid.service
systemctl start squid.service
echo "Done"| Capture domains from downloaded public blocklists (see SOURCES) and unify them in a single file. | Captura los dominios de las listas de bloqueo públicas descargadas (ver SOURCES) y las unifica en un solo archivo. |
Remove overlapping domains ('.sub.example.com' is a subdomain of '.example.com'), does homologation to Squid-Cache format and excludes false positives (google, hotmail, yahoo, etc.) with an allowlist (debugwl.txt).
|
Elimina dominios superpuestos ('.sub.example.com' es un dominio de '.example.com'), hace la homologación al formato de Squid-Cache y excluye falsos positivos (google, hotmail, yahoo, etc.) con una lista de permitidos (debugwl.txt).
|
Input:
com
.com
.domain.com
domain.com
0.0.0.0 domain.com
127.0.0.1 domain.com
::1 domain.com
domain.com.co
foo.bar.subdomain.domain.com
.subdomain.domain.com.co
www.domain.com
www.foo.bar.subdomain.domain.com
domain.co.uk
xxx.foo.bar.subdomain.domain.co.ukOutput:
.domain.com
.domain.com.co
.domain.co.uk| Remove domains with invalid TLDs (with a list of Public and Private Suffix TLDs: ccTLD, ccSLD, sTLD, uTLD, gSLD, gTLD, eTLD, etc., up to 4th level 4LDs). | Elimina dominios con TLD inválidos (con una lista de TLDs Public and Private Suffix: ccTLD, ccSLD, sTLD, uTLD, gSLD, gTLD, eTLD, etc., hasta 4to nivel 4LDs). |
Input:
.domain.exe
.domain.com
.domain.edu.coOutput:
.domain.com
.domain.edu.co|
Removes hostnames longer than 63 characters, as defined in RFC 1035, and other characters not admitted by IDN. It also converts domains with international, non-ASCII characters, used for homograph attacks, to the Punycode/IDNA format. |
Elimina hostnames de más de 63 caracteres, según define el RFC 1035, y otros caracteres no admitidos por IDN. También convierte los dominios con caracteres internacionales, no ASCII, usados para ataques homográficos, al formato Punycode/IDNA. |
Input:
bücher.com
café.fr
españa.com
köln-düsseldorfer-rhein-main.de
mañana.com
mūsųlaikas.lt
sendesık.com
президент.рфOutput:
xn--bcher-kva.com
xn--caf-dma.fr
xn--d1abbgf6aiiy.xn--p1ai
xn--espaa-rta.com
xn--kln-dsseldorfer-rhein-main-cvc6o.de
xn--maana-pta.com
xn--mslaikas-qzb5f.lt
xn--sendesk-wfb.com
Removes entries with invalid encoding, non-printable characters, whitespace, disallowed symbols, and any content that does not conform to the strict ASCII format for valid domain names (CP1252, ISO-8859-1, corrupted UTF-8, etc.). Converts the output to plain text with charset=us-ascii, ensuring a clean, standardized list ready for validation, comparison, or DNS resolution:
|
Elimina entradas con codificación incorrecta, caracteres no imprimibles, espacios en blanco, símbolos no permitidos y cualquier contenido que no se ajuste al formato ASCII estricto para nombres de dominio válidos (CP1252, ISO-8859-1, UTF-8 corrupto, etc.) y convierte la salida a texto sin formato charset=us-ascii, lo que garantiza una lista limpia y estandarizada, lista para validación, comparación o resolución DNS:
|
Input:
M-C$
-$
.$
0$
1$
23andmê.com
.òutlook.com
.ălibăbă.com
.ămăzon.com
.ăvăst.com
.amùazon.com
.aməzon.com
.avalón.com
.bĺnance.com
.bitdẹfender.com
.blóckchain.site
.blockchaiǹ.com
.cashpluÈ™.com
.dẹll.com
.diócesisdebarinas.org
.disnẹylandparis.com
.ebăy.com
.əməzon.com
.evo-bancó.com
.goglÄ™.com
.gooÄŸle.com
.googļę.com
.googlÉ™.com
.google.com
.ibẹria.com
.imgúr.com
.lloydÅŸbank.com
.mýetherwallet.com
.mrgreęn.com
.myẹthẹrwallet.com
.myẹthernwallet.com
.myethẹrnwallet.com
.myetheá¹™wallet.com
.myethernwallẹt.com
.nętflix.com
.paxfùll.com
.türkiyeisbankasi.com
.třezor.com
.westernúnion.com
.yòutube.com
.yăhoo.com
.yoütübe.co
.yoütübe.com
.yoütu.beOutput:
.google.com
Most of the SOURCES contain millions of invalid or nonexistent domains, so each domain is double-checked via DNS (in 2 steps) to exclude those entries from Blackweb. This process is performed in parallel and can be resource-intensive, depending on your hardware and network conditions. You can control concurrency with the PROCS variable:
|
La mayoría de las SOURCES contienen millones de dominios inválidos o inexistentes, por lo que cada dominio se verifica mediante DNS (en dos pasos) para excluir esas entradas de Blackweb. Este proceso se realiza en paralelo y puede consumir muchos recursos, dependiendo del hardware y las condiciones de la red. Puede controlar la concurrencia con la variable PROCS:
|
PROCS=$(($(nproc))) # Conservative (network-friendly)
PROCS=$(($(nproc) * 2)) # Balanced
PROCS=$(($(nproc) * 4)) # Aggressive (default)
PROCS=$(($(nproc) * 8)) # Extreme (8 or higher, use with caution)| For example, on a system with a Core i5 CPU (4 physical cores / 8 threads with Hyper-Threading): | Por ejemplo, en un sistema con una CPU Core i5 (4 núcleos físicos/8 subprocesos con Hyper-Threading): |
nproc → 8
PROCS=$((8 * 4)) → 32 parallel queries
PROCS values increase DNS resolution speed but may saturate your CPU or bandwidth, especially on limited networks like satellite links. Adjust accordingly. Real-time processing example:
|
PROCS aumentan la velocidad de resolución del DNS, pero pueden saturar la CPU o el ancho de banda, especialmente en redes limitadas como enlaces satelitales. Ajuste el sistema según corresponda. Ejemplo de procesamiento en tiempo real:
|
Processed: 2463489 / 7244989 (34.00%)Output (manual run):
HIT google.com
google.com has address 142.251.35.238
google.com has IPv6 address 2607:f8b0:4008:80b::200e
google.com mail is handled by 10 smtp.google.com.
FAULT testfaultdomain.com
Host testfaultdomain.com not found: 3(NXDOMAIN)| Remove government domains (.gov) and other related TLDs from BlackWeb. | Elimina de BlackWeb los dominios de gobierno (.gov) y otros TLD relacionados. |
Input:
.argentina.gob.ar
.mydomain.com
.gob.mx
.gov.uk
.navy.milOutput:
.mydomain.com
Run Squid-Cache with BlackWeb and any error sends it to SquidErrors.txt.
|
Corre Squid-Cache con BlackWeb y cualquier error lo envía a SquidErrors.txt.
|
Both bwupdate.sh and checksources.sh generate a log file (bwupdate.log / checksources.log) in the same directory as the script itself.
|
bwupdate.sh y checksources.sh generan un archivo de log (bwupdate.log / checksources.log) en el mismo directorio donde reside el script.
|
| Tag | Shows | English | Español |
|---|---|---|---|
SAVED: |
File name | The transfer completed and the file was written | La transferencia terminó completa y el archivo quedó escrito |
PARTIAL: |
Full URL | The download started and was cut off before finishing | La descarga arrancó y se cortó antes de terminar |
BUSY: |
Full URL | The server answered 5xx: it is up but not serving the list right now | El servidor respondió 5xx: está activo pero no sirve la lista en ese momento |
TIMEOUT: |
Full URL | The server did not answer at all | El servidor no respondió nada |
BROKEN: |
Full URL | The server answered 404 or 410: broken or nonexistent URL | El servidor respondió 404 o 410: URL rota o inexistente |
|
|
- ABPindo - indonesianadblockrules
- abuse.ch - hostfile
- Adaway - host
- adblockplus - advblock Russian
- adblockplus - antiadblockfilters
- adblockplus - easylistchina
- adblockplus - easylistlithuania
- anudeepND - adservers
- anudeepND - coinminer
- AssoEchap - stalkerware-indicators
- azet12 - KADhosts
- BarbBlock - blacklists
- BBcan177 - minerchk
- BBcan177 - MS-2
- BBcan177 - referrer-spam-blacklist
- betterwebleon - slovenian-list
- bigdargon - hostsVN
- BlackJack8 - iOSAdblockList
- BlackJack8 - webannoyances
- blocklistproject - everything
- chadmayfield - porn top
- chadmayfield - porn_all
- chainapsis - phishing-block-list
- cjx82630 - Chinese CJX's Annoyance List
- cobaltdisco - Google-Chinese-Results-Blocklist
- CriticalPathSecurity - Public-Intelligence-Feeds
- DandelionSprout - adfilt
- Dawsey21 - adblock-list
- Dawsey21 - main-blacklist
- developerdan - ads-and-tracking-extended
- Disconnect.me - simple_ad
- Disconnect.me - simple_malvertising
- Disconnect.me - simple_tracking
- Eallion - uBlacklist
- EasyList - EasyListHebrew
- ethanr - dns-blacklists
- fabriziosalmi - blacklists
- firebog - AdguardDNS
- firebog - Admiral
- firebog - Easylist
- firebog - Easyprivacy
- firebog - Kowabit
- firebog - neohostsbasic
- firebog - Prigent-Ads
- firebog - Prigent-Crypto
- firebog - Prigent-Malware
- firebog - RPiList-Malware
- firebog - RPiList-Phishing
- firebog - WaLLy3K
- frogeye - firstparty-trackers-hosts
- gardar - Icelandic ABP List
- greatis - Anti-WebMiner
- hagezi - dns-blocklists
- hexxium - threat-list/
- hoshsadiq - adblock-nocoin-list
- jawz101 - potentialTrackers
- jdlingyu - ad-wars
- kaabir - AdBlock_Hosts
- kevle1 - Windows-Telemetry-Blocklist - xiaomiblock
- liamja - Prebake Filter Obtrusive Cookie Notices
- malware-filter - URLhaus Malicious URL Blocklist
- malware-filter.- phishing-filter-hosts
- Matomo-org - spammers
- MBThreatIntel - malspam
- mine.nu - hosts0
- mitchellkrogza - Badd-Boyz-Hosts
- mitchellkrogza - hacked-domains
- mitchellkrogza - nginx-ultimate-bad-bot-blocker
- mitchellkrogza - strip_domains
- molinero - hBlock
- NanoAdblocker - NanoFilters
- neodevpro - neodevhost
- notracking - hosts-blocklists
- Oleksiig - Squid-BlackList
- openphish - feed
- pengelana - domains blocklist
- phishing.army - phishing_army_blocklist_extended
- piperun - iploggerfilter
- quidsup - notrack-blocklists
- quidsup - notrack-malware
- reddestdream - MinimalHostsBlocker
- RooneyMcNibNug - pihole-stuff
- Rpsl - adblock-leadgenerator-list
- ruvelro - Halt-and-Block-Mining
- ryanbr - fanboy-adblock
- scamaNet - blocklist
- ShadowWhisperer - Adult
- ShadowWhisperer - Cryptocurrency
- ShadowWhisperer - Dating
- ShadowWhisperer - Gambling
- ShadowWhisperer - Malware
- ShadowWhisperer - Risk
- ShadowWhisperer - Scam
- ShadowWhisperer - Shock
- ShadowWhisperer - Tracking
- ShadowWhisperer - Typo
- ShadowWhisperer - UrlShortener
- simeononsecurity/System-Wide-Windows-Ad-Blocker
- Someonewhocares - hosts
- stanev.org - Bulgarian adblock list
- StevenBlack - add.2o7Net
- StevenBlack - add.Risk
- StevenBlack - fakenews-gambling-porn-social
- StevenBlack - hosts
- StevenBlack - spam
- StevenBlack - uncheckyAds
- Stopforumspam - Toxic Domains
- sumatipru - squid-blacklist
- tomasko126 - Easylist Czech and Slovak filter list
- txthinking - blackwhite
- txthinking - bypass china domains
- Ultimate Hosts Blacklist - hosts
- Université Toulouse 1 Capitole - Blacklists UT1 - Olbat
- Université Toulouse 1 Capitole - Blacklists UT1
- Winhelp2002 - hosts
- yourduskquibbles - Web Annoyances Ultralist
- yous - YousList
- yoyo - Peter Lowe's Ad and tracking server list
- zoso - Romanian Adblock List
- google supported domains
- iana
- ipv6-hosts (Partial)
- publicsuffix
- Ransomware Database
- University Domains and Names Data List
- whoisxmlapi
BlackWeb's core is the domain/TLD blocklist (blackweb.txt, bwupdate.sh). The following are related subprojects, each with its own tooling and update flow.
|
El núcleo de BlackWeb es la lista de bloqueo de dominios/TLDs (blackweb.txt, bwupdate.sh). Lo siguiente son subproyectos relacionados, cada uno con sus propias herramientas y flujo de actualización.
|
fpack/ holds ACLs that are optional and not part of BlackWeb's core flow — bwupdate.sh never calls it. If interested, download the lists, run fpack.sh, and apply the Squid rules below yourself: ransomware file extensions, malicious User-Agents, and static web3 lists. fpack.sh generates the first two; web3 lists are curated by hand and not touched by the script.
|
fpack/ contiene ACLs opcionales, fuera del flujo core de BlackWeb — bwupdate.sh nunca lo invoca. Si te interesa, descarga las listas, corre fpack.sh, y aplica tú mismo las reglas de Squid de abajo: extensiones de archivo de ransomware, User-Agents maliciosos y listas web3 estáticas. fpack.sh genera las dos primeras; las listas web3 se curan a mano y el script no las toca.
|
fpack/
├── fpack.sh # generates rw/rwext.txt and ua/blockua.txt
├── rw/
│ ├── rw.txt # administrator ransomware blacklist (source)
│ ├── rwext.txt # Squid url_regex ACL (generated)
│ └── wl.txt # administrator ransomware whitelist (source)
├── ua/
│ └── blockua.txt # Squid browser ACL for bad User-Agents (generated)
└── web3/
├── web3domains.txt # Squid dstdomain ACL, web3/wallet domains (static)
└── web3tld.txt # Squid dstdomain ACL, web3/wallet TLDs (static)
wget,grep,sed,coreutils,util-linux
apt install -y wget grep sed coreutils util-linux
Either clone the whole BlackWeb repo, or download just the fpack/ folder with gitfolder.py (same tool vault subprojects use):
|
Clone todo el repo de BlackWeb, o descargue solo la carpeta fpack/ con gitfolder.py (la misma herramienta que usan los subproyectos de vault):
|
# Option 1: full clone
git clone --depth=1 https://github.com/maravento/blackweb.git
cd blackweb/fpack
# Option 2: folder-only download
wget -qO gitfolder.py https://raw.githubusercontent.com/maravento/vault/master/scripts/python/gitfolder.py
chmod +x gitfolder.py
python3 gitfolder.py https://github.com/maravento/blackweb/fpack
cd fpack
fpack.sh only produces Squid ACLs. Run it manually, as a non-root user — it is not scheduled by any installer:
|
fpack.sh solo produce ACLs para Squid. Se ejecuta manualmente, como usuario no root — ningún instalador la programa:
|
bash fpack.sh
Log: fpack.log, generated in fpack/, emptied at the start of every run.
|
Log: fpack.log, generado en fpack/, vaciado al inicio de cada ejecución.
|
# INSERT YOUR OWN RULE(S) HERE TO ALLOW ACCESS FROM YOUR CLIENTS
# Block Rule for Ransomware Extensions/Patterns (Optional)
# https://raw.githubusercontent.com/maravento/blackweb/master/fpack/rw/rwext.txt
acl block_ransomware url_regex -i "/path_to/rwext.txt"
http_access deny block_ransomware
# Block Rule for web3 domains (Optional)
# https://raw.githubusercontent.com/maravento/blackweb/master/fpack/web3/web3domains.txt
acl web3domains dstdomain "/path_to/web3domains.txt"
http_access deny web3domains
# Block Rule for web3 TLDs (Optional)
# https://raw.githubusercontent.com/maravento/blackweb/master/fpack/web3/web3tld.txt
acl web3tld dstdomain "/path_to/web3tld.txt"
http_access deny web3tld
# Block Rule for User-Agents (Optional)
# https://raw.githubusercontent.com/maravento/blackweb/master/fpack/ua/blockua.txt
acl bad_useragents browser -i "/path_to/blockua.txt"
http_access deny bad_useragents
Domain Filtering — removes overlapping domains, validates TLDs, and checks domain existence via DNS. It's a mandatory dependency of bwupdate.sh (called internally via domfilter.py; the update aborts if it fails) — but it can also be used as an independent project to process domain lists unrelated to BlackWeb:
|
Filtrado de Dominios — elimina dominios superpuestos, valida TLDs y verifica existencia de dominios vía DNS. Es una dependencia obligatoria de bwupdate.sh (se invoca internamente vía domfilter.py; la actualización aborta si falla) — pero también puede usarse como proyecto independiente para procesar listas de dominios ajenas a BlackWeb:
|
python3 domfilter.py --input mylst.txtdofi/
├── domcheck.sh # Checks domain existence with the host command
└── domfilter.py # Removes overlapping domains, validates TLDs
- Python 3.12.3, Bash 5.2.21
python3-requests(required bydomfilter.py),bind9-host,findutils,coreutils,util-linux(required bydomcheck.sh)
apt install -y python3 python3-requests bind9-host findutils coreutils util-linux# Option 1: full clone
git clone --depth=1 https://github.com/maravento/blackweb.git
cd blackweb/dofi
# Option 2: folder-only download
wget -qO gitfolder.py https://raw.githubusercontent.com/maravento/vault/master/scripts/python/gitfolder.py
chmod +x gitfolder.py
python3 gitfolder.py https://github.com/maravento/blackweb/dofi
cd dofi
Ensure the input list has no http://, https://, or www. prefixes. What it does:- Downloads public suffix TLDs from multiple sources. - Removes invalid or duplicate TLDs. - Filters domains to ensure they end with a valid TLD. - Removes overlapping domains. - Excludes duplicates from previously validated domains. - Outputs results to a file. |
Asegúrese de que la lista de entrada no tenga prefijos http://, https:// o www.. Qué hace:- Descarga TLD de sufijo público de varias fuentes. - Elimina TLD no válidos o duplicados. - Filtra dominios para garantizar que terminen con un TLD válido. - Elimina dominios superpuestos. - Excluye duplicados de dominios previamente validados. - Envía los resultados a un archivo. |
python3 domfilter.py --input mylst.txt
Replace mylst.txt with the name of your domain list. By default, output goes to output.txt and removed lines to removed.txt — customize with:
|
Reemplace mylst.txt con el nombre de su lista de dominios. Por defecto, la salida va a output.txt y las líneas eliminadas a removed.txt — personalice con:
|
python3 domfilter.py --input mylst.txt --output outlst.txt --removed removelst.txtTLD Includes: ccTLDs, gTLDs, sTLDs, eTLDs, and 4LDs (file: tlds.txt)
Checks if each domain exists, using the host command, and cleans the input list, separating the output as follows:- hit.txt: existing domains from your list.- fault.txt: non-existent domains removed.
|
Verifica si cada dominio existe, con el comando host, y limpia la lista de entrada, separando la salida de la siguiente manera:- hit.txt: dominios existentes de su lista.- fault.txt: dominios inexistentes eliminados.
|
chmod +x domcheck.sh
bash domcheck.sh my_domain_list.txt
Replace my_domain_list.txt with your input list. Optional: replace 50 with the number of parallel processes (default: nproc × 4):
|
Reemplace my_domain_list.txt con su lista de entrada. Opcional: reemplace 50 con el número de procesos paralelos (por defecto: nproc × 4):
|
bash domcheck.sh my_domain_list.txt 502026-07-07 13:29:58 domcheck start...
2026-07-07 13:29:58 Step 1...
2026-07-07 13:30:17 OK
2026-07-07 13:30:17 Step 2...
2026-07-07 13:30:28 hit.txt: domains successfully resolved
2026-07-07 13:30:28 fault.txt: unresolved domains
2026-07-07 13:30:28 Summary:
2026-07-07 13:30:28 Input domains : 7896
2026-07-07 13:30:28 Resolved : 7349
2026-07-07 13:30:28 Unresolved : 547
2026-07-07 13:30:28 Elapsed time : 30s
2026-07-07 13:30:29 domcheck done at: mar 07 jul 2026 13:30:29 -05- Awesome Open Source: Blackweb
- Community IPfire: url filter and self updating blacklists
- covert.io: Getting Started with DGA Domain Detection Research
- Crazymax: WindowsSpyBlocker
- egirna: Allowing/Blocking Websites Using Squid
- Jason Trost: Getting Started with DGA Domain Detection Research
- Kandi Openweaver: Domains Blocklist for Squid-Cache
- Kerry Cordero: Blocklists of Suspected Malicious IPs and URLs
- Keystone Solutions: blocklists
- Lifars: Sites with blocklist of malicious IPs and URLs
- Opensourcelibs: Blackweb
- OSINT Framework: Domain Name/Domain Blacklists/Blackweb
- Osintbay: Blackweb
- Reddit: Blackweb
- Secrepo: Samples of Security Related Data
- Segu-Info: Análisis de malware y sitios web en tiempo real
- Segu-Info: Dominios/TLD dañinos que pueden ser bloqueados para evitar spam y #phishing
- Soficas: CiberSeguridad - Protección Activa
- Stackoverflow: Blacklist IP database
- Wikipedia: Blacklist_(computing)
- Xploitlab: Projects using WindowsSpyBlocker
- Zeltser: Free Blocklists of Suspected Malicious IPs and URLs
- Zenarmor: How-to-enable-web-filtering-on-OPNsense-proxy?
|
|
wget https://raw.githubusercontent.com/maravento/blackweb/refs/heads/master/bwupdate/tools/checksources.sh
chmod +x checksources.sh
./checksources.she.g:
[?] Enter domain to search: kickass.to
[*] Searching for 'kickass.to'...
[+] Domain found in: https://github.com/fabriziosalmi/blacklists/releases/download/latest/blacklist.txt
[+] Domain found in: https://hostsfile.org/Downloads/hosts.txt
[+] Domain found in: https://raw.githubusercontent.com/blocklistproject/Lists/master/everything.txt
[+] Domain found in: https://raw.githubusercontent.com/hagezi/dns-blocklists/main/wildcard/ultimate-onlydomains.txt
[+] Domain found in: https://raw.githubusercontent.com/Ultimate-Hosts-Blacklist/Ultimate.Hosts.Blacklist/master/hosts/hosts0
[+] Domain found in: https://sysctl.org/cameleon/hosts
[+] Domain found in: https://v.firebog.net/hosts/Kowabit.txt
DoneSpecial thanks to: Jhonatan Sneider
| This project uses a dual-licensing model to balance software freedom with content protection: | Este proyecto utiliza un modelo de licencia dual para equilibrar la libertad del software con la protección del contenido: |
| Content | Licensed Under |
|---|---|
| Scripts, Binaries, Infrastructure | |
| RAG, Workers, Specialized Modules, Docs |
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.
|
Due to recent arbitrary changes in computer terminology, it is necessary to clarify the meaning and connotation of the term blacklist, associated with this project:
In computing, a blacklist, denylist or blocklist is a basic access control mechanism that allows through all elements (email addresses, users, passwords, URLs, IP addresses, domain names, file hashes, etc.), except those explicitly mentioned. Those items on the list are denied access. The opposite is a whitelist, which means only items on the list are let through whatever gate is being used. Source Wikipedia Therefore, blacklist, blocklist, blackweb, blackip, whitelist and similar, are terms that have nothing to do with racial discrimination. |
Debido a los recientes cambios arbitrarios en la terminología informática, es necesario aclarar el significado y connotación del término blacklist, asociado a este proyecto:
En informática, una lista negra, lista de denegación o lista de bloqueo es un mecanismo básico de control de acceso que permite a través de todos los elementos (direcciones de correo electrónico, usuarios, contraseñas, URL, direcciones IP, nombres de dominio, hashes de archivos, etc.), excepto los mencionados explícitamente. Esos elementos en la lista tienen acceso denegado. Lo opuesto es una lista blanca, lo que significa que solo los elementos de la lista pueden pasar por cualquier puerta que se esté utilizando. Fuente Wikipedia Por tanto, blacklist, blocklist, blackweb, blackip, whitelist y similares, son términos que no tienen ninguna relación con la discriminación racial. |
