Saltar al contenido
HIDDENIO

Metodología

Cómo se construye el índice de HiddenIO: qué recopilamos, cómo se analizan y fusionan los registros, y qué significan (y qué no) nuestras etiquetas.

Collection

HiddenIO is a passive, source-driven index. We do not scan IP ranges and we never probe, connect to, or test the servers behind the configurations we list. The pipeline only fetches documents that are already public — public channels, websites, repositories, pastebins and raw lists — exactly as an ordinary visitor would.

Our crawler identifies itself as HiddenIOBot/1.0 in the User-Agent string, honors robots.txt, and keeps request rates per source well within courtesy limits. Sources are only added after human review (see the source policy), and any publisher can opt out — details are in the bot policy.

Parsing

Every fetched document is parsed by a protocol-aware parser (vless, vmess, trojan, Shadowsocks, ShadowsocksR, Hysteria/Hysteria2, TUIC, SOCKS, HTTP and VPN formats). Parsers are deliberately lenient: they accept the field-name and encoding variants seen in the wild (SIP002 and legacy Shadowsocks, URL-safe or unpadded Base64, common vmess JSON keys) and record a parse error rather than crashing on unknown syntax.

Normalization

Extracted entries are normalized before comparison: hostnames are lower-cased, default ports are made explicit, percent-encoding is decoded, and transport/security parameters are mapped to a canonical vocabulary (type → network, security → none/tls/reality/xtls).

Deduplication

The same configuration often appears in many sources with cosmetic differences. We deduplicate on three fingerprints:

  • Endpoint fingerprint — protocol + address + port, used for source-coverage counting.
  • Normalized fingerprint — endpoint plus normalized transport and security fields.
  • Semantic fingerprint — additionally includes credential fields (UUID/password), so entries that

differ only in remark name or parameter order collapse into one record, while different credentials stay separate.

A record's source_count and observation_count reflect how many distinct sources and fetches have shown it — a popularity signal, not a quality or safety signal.

Freshness

Each record carries a freshness bucket based on the time it was last observed in any source:

See the detailed freshness definitions for edge cases such as reappearances.

Country classification

Country labels are heuristic. We combine (in order) explicit country tokens in the entry name, flag emojis (e.g. 🇩🇪), ISO codes in hostnames, and — optionally — an offline IP-geolocation database for entries whose host resolves to a stable IP. All of these can be wrong: a server physically in one country can be named after another. Country labels are best-effort hints, never verified facts.

Limitations

The index describes what was published, nothing more. We cannot verify that any server is online, that it accepts connections, that it is trustworthy, or that using it is legal where you are. Entries can be stale, misnamed, duplicated under different fingerprints, or deliberately planted. Freshness reflects observation recency, not availability. Use judgment, report problems via the report form, and read the disclaimer.

Freshness buckets

Bucketfresh
DefinitionSeen in a source within the last 3 days
Bucketrecent
DefinitionLast seen 3–14 days ago
Bucketaging
DefinitionLast seen 14–45 days ago
Bucketstale
DefinitionLast seen 45–90 days ago
Bucketarchived
DefinitionNot seen for over 90 days; kept for reference and excluded from default views