The Protocol That Holds the Internet Together

Inside one organisation, OSPF finds the shortest path. Between organisations the question is which neighbour I am willing to send traffic through, and whose traffic I will carry. BGP (Border Gateway Protocol) answers that, and every route on the public internet reaches your ISP through it. The routing lesson introduced autonomous systems and the decision order; this one covers how BGP sessions run, why the internet's shape follows money, and how routes go wrong.

Autonomous Systems and Business Relationships

An autonomous system (AS) is a network with its own routing policy, identified by a number from a regional internet registry. The examples below use ASNs 64496–64511, which are reserved for documentation (like 203.0.113.0/24 for addresses).

Links between ASes come in two kinds. In transit, a customer pays a provider to carry its traffic to and from the entire internet. In peering, two networks exchange traffic for their own and their customers' addresses, usually without payment, often at internet exchange points (IXPs) where hundreds of networks share a switch.

Money explains the export rules every operator follows, the "valley-free" or Gao-Rexford model:

Route learned from Export to customers Export to peers Export to providers
a customer yes yes yes
a peer yes no no
a provider yes no no

Announcing customer routes everywhere earns money; passing one provider's routes to another would mean carrying, for free, traffic you pay for on both sides. Operators encode the same economics in LOCAL_PREF: customer routes highest, then peers, then transit.

Inside a BGP Session

BGP runs over TCP port 179 with four messages: OPEN (AS number, router ID, hold time, capabilities such as IPv6 and 4-byte ASNs), KEEPALIVE, UPDATE (announce prefixes with attributes, or withdraw them) and NOTIFICATION (a fatal error, then the session closes). Once Established, each side sends its full table (roughly a million IPv4 prefixes for a full feed), then only changes. The state names mislead: Active means the router is trying to connect and failing.

Every eBGP session needs three safeguards: an outbound prefix list so you only announce your own space, an inbound filter so a neighbour can only send what it is entitled to, and a max-prefix limit that drops the session if the neighbour suddenly sends far more routes than expected. FRRouting refuses to exchange eBGP routes on a session with no policy at all (RFC 8212 behaviour), so a new session can be Established with zero prefixes.

eBGP and iBGP

eBGP runs between ASes, normally over a direct link. iBGP distributes what the border routers learned to the rest of the same AS. Two differences cause most confusion:

  • iBGP does not add the AS number, so AS_PATH cannot catch loops inside the AS. Instead, a route learned from one iBGP peer is never passed to another, which forces a full mesh: n(n-1)/2 sessions, 190 for 20 routers. Large networks use route reflectors, which may re-advertise to their clients and tag routes with ORIGINATOR_ID and CLUSTER_LIST to stop loops.
  • NEXT_HOP is not changed on iBGP. Internal routers receive the external neighbour's address as next hop, and if the IGP has no route to it the BGP route is unusable. Carry the border links in the IGP or set next-hop-self on the border routers.

The Decision Process as Code

When several routes to a prefix survive filtering, the router compares attributes in a fixed order and the first difference decides. Python's tuple comparison works exactly that way:

from dataclasses import dataclass

@dataclass
class Route:
    name: str
    local_pref: int = 100
    as_path: tuple = ()
    origin: int = 0          # 0 = IGP, 1 = EGP, 2 = incomplete
    med: int = 0
    ebgp: bool = True
    igp_cost: int = 0
    router_id: str = "0.0.0.0"

LABELS = ["LOCAL_PREF", "AS_PATH length", "ORIGIN", "MED", "eBGP over iBGP",
          "IGP cost to next hop", "router ID"]

def key(r):
    return (-r.local_pref, len(r.as_path), r.origin, r.med, 0 if r.ebgp else 1,
            r.igp_cost, tuple(int(o) for o in r.router_id.split(".")))

def explain(routes):
    best, runner_up = sorted(routes, key=key)[:2]
    for label, a, b in zip(LABELS, key(best), key(runner_up)):
        if a != b:
            return f"best: {best.name}  (beats {runner_up.name} on {label})"

transit_a = Route("transit-A", as_path=(64510, 64499), router_id="10.0.0.1")
transit_b = Route("transit-B", as_path=(64511, 64505, 64499), router_id="10.0.0.2")
peer_ix   = Route("IX-peer",   as_path=(64499,), local_pref=200, router_id="10.0.0.3")
print(explain([transit_a, transit_b]))
print(explain([transit_a, transit_b, peer_ix]))

near_b = Route("transit-B", as_path=(64511, 64499), igp_cost=5, router_id="10.0.0.2")
far_a  = Route("transit-A", as_path=(64510, 64499), igp_cost=20, router_id="10.0.0.1")
print(explain([far_a, near_b]))
best: transit-A  (beats transit-B on AS_PATH length)
best: IX-peer  (beats transit-A on LOCAL_PREF)
best: transit-B  (beats transit-A on IGP cost to next hop)

What the model leaves out, and what surprises people on real routers:

  • MED is by default compared only between routes from the same neighbouring AS.
  • AS_PATH length ignores link speed. Two ASes over a congested ocean cable beat three next door. BGP is policy, not performance.
  • The IGP-cost step is hot-potato routing: each network hands traffic off at the exit nearest where it entered, so forward and return paths between two hosts often differ.

You control outbound traffic fully with LOCAL_PREF. Inbound traffic depends on other networks' choices, so you can only influence it: AS path prepending (repeating your ASN to make a path look longer), more-specific announcements on the preferred link, or communities, tags like 64501:80 whose meaning providers publish ("set local-pref 80", "do not announce to peer X"). The well-known BLACKHOLE community (65535:666) asks a provider to drop traffic to a prefix under DDoS attack.

Looking Up Real Routes

The RIPE RIS project collects BGP feeds from hundreds of networks, and RIPEstat serves them as a public JSON API. This script finds the prefix and origin AS covering an address and summarises the AS paths the collectors see:

import json
import sys
import urllib.parse
import urllib.request
from collections import Counter

def stat(call, **params):
    url = f"https://stat.ripe.net/data/{call}/data.json?" + urllib.parse.urlencode(params)
    with urllib.request.urlopen(url, timeout=30) as resp:
        return json.load(resp)["data"]

ip = sys.argv[1] if len(sys.argv) > 1 else "1.1.1.1"
overview = stat("prefix-overview", resource=ip)
prefix = overview["resource"]                      # the announced prefix covering the IP
for a in overview["asns"]:
    rpki = stat("rpki-validation", resource=a["asn"], prefix=prefix)
    print(f"{ip} in {prefix}, origin AS{a['asn']} ({a['holder']}), RPKI {rpki['status']}")

lg = stat("looking-glass", resource=prefix)
paths = [p["as_path"].split() for rrc in lg["rrcs"] for p in rrc["peers"]]
lengths = Counter(len(set(p)) for p in paths)      # set() ignores prepending
print(f"{len(paths)} paths from {len(lg['rrcs'])} collectors;",
      "unique ASes per path:", dict(sorted(lengths.items())))
1.1.1.1 in 1.1.1.0/24, origin AS13335 (CLOUDFLARENET - Cloudflare, Inc.), RPKI valid
360 paths from 23 collectors; unique ASes per path: {1: 1, 2: 300, 3: 49, 4: 10}

Most paths are two ASes long: the collector's peer, then Cloudflare. Large content networks peer directly with thousands of networks, flattening the old tier-1 hierarchy. For path problems on your own connections, mtr -z shows the AS of each hop (see network monitoring and troubleshooting).

Leaks and Hijacks

BGP believes neighbours unless they are filtered. Three failures recur:

  • Route leak: a network breaks the valley-free rules, typically a multihomed customer passing one provider's full table to another. Huge volumes of traffic then squeeze through the small network and are dropped.
  • Origin hijack: an AS announces a prefix it does not hold, and wins wherever its path looks shorter.
  • More-specific hijack: a longer prefix inside someone else's block wins everywhere it spreads, because longest prefix match beats every BGP attribute. In the 2008 YouTube incident, a provider announced a more-specific of YouTube's space to block it nationally; it leaked upstream and took YouTube offline for much of the internet for about two hours.

Most incidents are accidents. The defences are prefix filters built from IRR route objects (generated with tools like bgpq4), max-prefix limits, and RPKI.

RPKI Route Origin Validation

Address holders sign ROAs in their RIR's portal: "prefix P, up to length maxLength, may be originated by AS N". Validators such as Routinator or rpki-client verify them and feed routers, which mark each route valid, invalid or not-found:

import ipaddress

ROAS = [(ipaddress.ip_network("203.0.113.0/24"), 24, 64500),    # (prefix, maxLength, origin)
        (ipaddress.ip_network("198.51.100.0/22"), 24, 64501)]

def validate(prefix, origin):
    '''RFC 6811 route origin validation.'''
    net = ipaddress.ip_network(prefix)
    covering = [r for r in ROAS if net.version == r[0].version and net.subnet_of(r[0])]
    if not covering:
        return "not-found"      # no ROA: usually still accepted
    if any(origin == asn and net.prefixlen <= max_len for _, max_len, asn in covering):
        return "valid"
    return "invalid"            # covered, but wrong origin or too specific

for prefix, asn in [("203.0.113.0/24", 64500), ("203.0.113.0/24", 64509),
                    ("203.0.113.0/25", 64500), ("198.51.101.0/24", 64501),
                    ("192.0.2.0/24", 64502)]:
    print(f"{prefix:<18} AS{asn:<6} {validate(prefix, asn)}")
203.0.113.0/24     AS64500  valid
203.0.113.0/24     AS64509  invalid
203.0.113.0/25     AS64500  invalid
198.51.101.0/24    AS64501  valid
192.0.2.0/24       AS64502  not-found

The third line is the trap: a /25 from the legitimate origin is invalid because it exceeds maxLength. Practical consequences:

  • Cover everything you announce, including more-specifics used during DDoS mitigation, or your own routes get dropped.
  • Keep maxLength tight. A loose maxLength lets a hijacker announce a more-specific with your ASN forged as origin and still be "valid".
  • ROV checks only the origin. A forged path ending in the right ASN passes. ASPA, where each AS publishes its authorised providers so leaks and forged paths become detectable, is the emerging fix.

Troubleshooting

Symptom Likely cause
Stuck in Active/Connect TCP 179 blocked, wrong address or remote-as, or eBGP neighbour not directly connected (TTL 1)
Established, 0 prefixes inbound policy rejects everything, or no policy (RFC 8212)
Prefix not visible on the internet prefix not in your routing table (a network statement needs it), outbound filter, no IRR object at the upstream, RPKI invalid, or longer than /24 (IPv4) or /48 (IPv6), which most networks filter
Session flaps every few minutes hold timer expiring: CPU starvation, or an MTU problem stalling large UPDATEs while small keepalives pass

Practice

Run the lookup script on your own public IP to find your ISP's AS and ROA status. Then change key() so MED is compared only between routes whose first AS is the same, and build a case where the winner changes.