Akamai Bot Manager: Sensor Data and `_abck`
The Cookie Records the Verdict, Not the Evidence
Every request to an edge protected by a sensor system carries a cookie whose value looks like punctuation. That value is not the sensor data. It is the score the edge computed from a sensor payload your browser posted earlier, plus a handful of segments the server rewrites on every response. The evidence lives in the POST body; the cookie is what the next request is judged on. The score that cookie feeds is the same broad one described in How Anti-Bots Work, and the payload that produces it is built from the same browser surface as Browser Fingerprinting.
This lesson works through a synthetic payload of the shape these systems use: a delimited text body, a fixed field vocabulary, a length prefix on every field, and a score carried forward into the next request. Nothing here is captured from a live site. The value is in seeing the structure, the byte accounting and the failure modes clearly enough that you can recognise them in a capture you are authorised to analyse, and in reading them through the cookie lifecycle that Proof-of-Work and Clearance Cookies covers from the clearance side. When a session stops being accepted, the capture that shows why is usually read with the method in Challenge Debugging with Recorded Traffic.
Where the Payload Comes From and What It Costs to Collect
The sensor script runs on the page and gathers fields into categories: index counters, session identifiers, platform reads, browser capability reads, timing deltas, behaviour samples, and a small set of shaping probes. The category table in the run below prices each one by what the site pays to collect it, and the ordering is the interesting part.
Platform reads are thirteen fields and thirty percent of the body, at no extra network cost. Browser capability reads are twelve fields and thirty-two percent, and the top of that column says why: a canvas or WebGL readback forces a GPU round trip that a script which cannot do it gives away instantly. Behaviour samples are four fields and two percent of the bytes, and are by far the most expensive row to store and to review on the defender's side. Reading that table as a cost sheet rather than a feature list is the useful move: it tells you which fields a site has a reason to keep asking for every time, and which it can afford to sample.
The Body Is a Delimited Text Record, and the Codec Is the Cheap Half
The transport is a plain POST with content-type: text/plain to a path of the
form /akam/{build}/{sensor id}/sensor_data. The body is a flat sequence of
-field-value- records, with each value preceded by a four-digit zero-padded
length.
The codec in the snippet encodes that structure and asserts a byte-identical round trip, which is the only requirement worth stating up front: the edge recomputes a record hash over the body it received, so reordering two fields, shortening one length prefix, or sending an unchanged value twice are all detectable. None of them change the structured payload you believe you are sending, which is exactly why they survive review. The codec is a few dozen lines and it is the easy part; treat it as solved and spend your attention on the vocabulary it carries.
The Vocabulary Is the Half That Has to Be Re-Derived Every Time
Forty-two fields is not a large surface, and the origin table shows how
concentrated it is: thirty-seven distinct browser surfaces feed them, and a
third of them come from three sources. aj_indx and its companions are counters
that exist to detect a loop that posts the same record twice. session_id and
session_key tie the payload to the cookie jar. ab_score is the previous
verdict, carried forward so the edge can weight its own history.
The fragility table is the part to keep. The deployed script for one build wants ten fields, the next build wants fifteen, and two of those are absent from a payload built against the older vocabulary. Such a payload is not rejected on syntax. It is scored on absence, and absence is recorded in the cookie. That is the asymmetry that makes this class of system hard to reason about from the outside: the failure is not an error, it is a slightly worse score, and a worse score usually only manifests two requests later.
The _abck Lifecycle Has Three Stages and One Trap
The lifecycle table walks one session. Stage one is a single token, the pre-instrument marker, set by a document response before any payload exists. Stage two is three segments after a sensor post: the hardware segment, the mode and record-length segment, and the response hash. Stage three is a delimited outbound segment that the response rewrites on every document.
The trap is in the last row. A stale replay leaves the cookie flagged rather than clearing it, so re-presenting a flagged value keeps the identity flagged no matter how many clean payloads follow. Once a session reaches the abnormal segment, the useful operation is to abandon the cookie and the profile together and start a new identity, not to keep posting. Reading the value as a state machine with three stages, rather than as an opaque string, is what makes the "flagged forever" behaviour obvious instead of mysterious.
'''Inventory a synthetic Akamai Bot Manager sensor payload.
The field values below are invented for this lesson. The payload SHAPE is what
real captures show: a plain-text body of name/length/value records joined by a
five-character framing delimiter, posted to a sensor endpoint, with the verdict
written back into a Set-Cookie header.
The length field here is zero-padded to four digits so a length-driven reader
is unambiguous. Captures vary in that detail; nothing else in this file touches
the network, so the output is identical on every run.
'''
DELIM = "-1_1-"
# name -> (category, origin, value). Origins are the browser surfaces a real
# sensor script reads; categories are how a defender budgets the collection.
RAW_FIELDS = {
"aj_indx": ("index", "counter", "1491928374"),
"aj_typ": ("index", "counter", "UNKNOWN"),
"aj_indx_2": ("index", "counter", "0"),
"session_id": ("session", "cookie and storage", "7f3c9a11e2b84d05a6c7"),
"session_key": ("session", "storage probe", "4b1f0d9c8a7e6f5d"),
"ab_score": ("verdict", "previous sensor verdict", "0.42"),
"platform": ("platform", "navigator.userAgentData", "Win32"),
"platformVersion": ("platform", "navigator.userAgentData", "15.0.0"),
"hardware": ("platform", "navigator.hardwareConcurrency", "12"),
"cpu": ("platform", "navigator", "AMD Ryzen 7 5700G"),
"browser": ("platform", "userAgent parse", "Chrome"),
"browser_v": ("platform", "userAgent parse", "130.0.0.0"),
"moz": ("browser", "canvas 2d hash", "1743902"),
"dnt": ("browser", "navigator.doNotTrack", "unspecified"),
"fp": ("browser", "canvas and webgl hash", "9481273.5"),
"websec_app_id": ("browser", "document.referrer host", "shop.example"),
"websec_bc": ("browser", "cookie write test", "1"),
"websec_frame": ("browser", "frame ancestry", "0"),
"tz": ("platform", "Intl.timeZone", "America/Chicago"),
"tz_off": ("platform", "Date offset", "18000"),
"lang": ("platform", "navigator.language", "en-US"),
"screen": ("platform", "screen dimensions", "1920x1080x24"),
"sr": ("platform", "screen dimensions", "1920x1080"),
"ds": ("platform", "device pixel ratio", "1"),
"rqst": ("timing", "performance timeline", "18|41|60|63|71"),
"rqdd": ("timing", "request start deltas", "12|18|21|24|26"),
"rqurl": ("timing", "referrer and path", "/checkout/step2"),
"m_bdr": ("behaviour", "pointer sample count", "0"),
"m_ti": ("behaviour", "pointer sample ticks", "0"),
"m_ul": ("behaviour", "longest pointer delta", "0"),
"m_dm": ("behaviour", "key, wheel and touch", "0|0|0"),
"rqet": ("timing", "request end timestamp", "1767225600123"),
"bcv": ("browser", "canvas vector crosscheck", "3|0|1"),
"wd": ("browser", "window and doc sizes", "1920|947|1920|0|0|24|1"),
"cloc": ("browser", "geolocation result", "denied"),
"ckey": ("browser", "document.cookie probe", "8c1f0a2d"),
"fbe": ("browser", "storage enumeration", "1|2|0|1|0"),
"r_dat": ("telemetry", "response shaping check", "0"),
"r_num": ("telemetry", "numeric sanity probe", "1"),
"r_pos": ("telemetry", "position fields present", "1"),
"sw": ("platform", "software inventory", "12|2|1"),
"ch": ("browser", "UA client hint model", "UA-CH present"),
}
# What the deployed sensor script for a given build is known to emit.
REQUIRED_BY_VERSION = {
"build 3.x": ["aj_indx", "session_id", "ab_score", "platform", "fp",
"rqst", "rqdd", "websec_app_id", "m_dm", "ckey"],
"build 4.x": ["aj_indx", "session_id", "ab_score", "platform", "platformVersion",
"fp", "rqst", "rqdd", "websec_app_id", "m_dm", "ckey",
"fp_2", "sr", "tz_off", "audio_hash"],
}
PRESENT = set(RAW_FIELDS)
LENGTH_WIDTH = 4
def encode(fields):
'''Build the wire body: delimiter, name, four-digit length, then the value.'''
out = []
for name, (_category, _origin, value) in fields.items():
out.append("{}{}-{:0{}d}{}".format(
DELIM, name, len(value), LENGTH_WIDTH, value))
return "".join(out)
def decode(wire):
'''Parse the wire body back, driven by the declared lengths, not by guesses.
Returns (records, problems). A length-driven read needs no field table,
which is why this half is cheap; the field table is the other half, and it
is the half that goes stale.
'''
width, size = len(DELIM), LENGTH_WIDTH
records, problems = [], []
pos = 0
if wire[:width] != DELIM:
raise ValueError("body does not start with the framing delimiter")
pos = width
while pos < len(wire):
dash = wire.index("-", pos)
name = wire[pos:dash]
pos = dash + 1
length = int(wire[pos:pos + size])
pos += size
records.append((name, wire[pos:pos + length]))
pos += length
if pos >= len(wire):
break
if wire[pos:pos + width] != DELIM:
raise ValueError("framing lost after field " + name)
pos += width
return records, problems
wire = encode(RAW_FIELDS)
parsed, parse_problems = decode(wire)
assert len(parsed) == len(RAW_FIELDS), "round trip lost a field"
assert dict(parsed) == {k: v[2] for k, v in RAW_FIELDS.items()}, "round trip mismatch"
print("synthetic sensor payload: {} fields, {:,} bytes in the body".format(
len(parsed), len(wire)))
print("round trip through the codec: byte identical, 0 parse problems")
print("endpoint shape: POST /akam/<build>/<sensor id>/sensor_data, content-type text/plain")
print()
print("{:<20} {:<11} {:>5} {:<27} {}".format(
"field", "category", "bytes", "origin", "value"))
print("-" * 104)
by_category, by_origin = {}, {}
for name, value in parsed:
category, origin, _ = RAW_FIELDS[name]
by_category[category] = by_category.get(category, 0) + len(value)
by_origin[origin] = by_origin.get(origin, 0) + 1
print("{:<20} {:<11} {:>5} {:<27} {}".format(
name, category, len(value), origin, value[:24]))
total = sum(by_category.values())
print()
print("{:<12} {:>7} {:>8} {:>8} {}".format(
"category", "fields", "bytes", "share", "what a defender pays for it"))
print("-" * 100)
SPEND = {
"index": "loopback counters, free to compute",
"session": "cookie and storage reads, writes one cookie",
"verdict": "one field, but it is the one the next request is scored on",
"platform": "navigator and Intl reads, no extra round trip",
"browser": "canvas and webgl readback: forces a GPU round trip",
"timing": "performance timeline reads, free but they leak interaction shape",
"behaviour": "event listeners, the most expensive row to store and to review",
"telemetry": "shaping probes that exist to detect a fake navigator",
}
for category in sorted(by_category, key=lambda c: -by_category[c]):
count = sum(1 for n, _ in parsed if RAW_FIELDS[n][0] == category)
print("{:<12} {:>7} {:>8} {:>7.1f}% {}".format(
category, count, by_category[category],
100.0 * by_category[category] / total, SPEND[category]))
print()
print("origin concentration: {} distinct browser surfaces feed these fields".format(len(by_origin)))
for origin, count in sorted(by_origin.items(), key=lambda kv: (-kv[1], kv[0]))[:6]:
print(" {:<28} {} field(s)".format(origin, count))
print()
print("version fragility: what the deployed script for each build is known to emit")
print("{:<10} {:>7} {:>6} {:>8} {}".format(
"build", "wanted", "sent", "missing", "fields this capture does not carry"))
print("-" * 100)
for version, required in REQUIRED_BY_VERSION.items():
missing = [f for f in required if f not in PRESENT]
print("{:<10} {:>7} {:>6} {:>8} {}".format(
version, len(required), len(required) - len(missing), len(missing),
", ".join(missing) if missing else "none, complete for this build"))
print()
print("the 4.x list wants fields this capture does not have. A payload missing them")
print("is not rejected on syntax; it is scored on absence, and the score is what the")
print("_abck cookie records. The codec is the cheap half. The vocabulary is the half")
print("that has to be re-derived every time the vendor rotates the script.")
print()
print("_abck lifecycle: the shapes one cookie value takes across a session")
COOKIE_SCRIPT = [
("first document", "~hw",
"no payload posted yet: the pre-instrument marker"),
("first sensor post", "~-hw~m~4~3b1c9f0a2d~mr~8f21e7d4c0~",
"hardware segment present, status segment present"),
("flagged post", "~-hw~m~4~c0ffee11dead~mr~000000000000~-ab~1~",
"abnormal marker: scored, and scored badly"),
("next document", "~-hw~m~4~3b1c9f0a2d~mr~8f21e7d4c0~-ou~ok~0~mr~8f21e7d4c0~",
"punctuated status form: ok, plus an outbound segment the response refreshed"),
("stale replay", "~-hw~m~4~c0ffee11dead~mr~000000000000~-ou~ir~0~mr~000000000000~",
"re-presented unchanged: the edge leaves it flagged rather than clearing it"),
]
for stage, value, note in COOKIE_SCRIPT:
print(" {:<18} {}".format(stage, value))
print(" {:<18} {}".format("", note))
print()
print("stage 1 is one token. Stage 2 is three segments. Stage 3 is a delimited")
print("outbound segment that the response rewrites on every document. The lifecycle")
print("is: no payload, then a clean status, then a punctuated refresh per request.")
print()
print("segment interpretation, as read from captures. The vocabulary is not")
print("documented by the vendor and it does change between builds.")
print("{:<9} {:<19} {:<30} {}".format("segment", "where it sits", "reading", "what to do"))
print("-" * 104)
SEGMENTS = [
("~hw", "whole value", "pre-instrument marker", "a sensor_data POST should follow"),
("-hw-", "delimited", "hardware segment", "only valid after a posted payload"),
("~m~n~", "inside -hw-", "mode and record length", "recheck the posted body size"),
("~mr~", "end of -hw-", "response hash", "changes per request; a repeat is odd"),
("-ou-", "delimited", "outbound segment", "refreshed by the response, not by you"),
("~ok~", "inside -ou-", "status: accepted", "session clean, carry on"),
("~ir~", "inside -ou-", "status: invalid record", "payload rejected, re-post from scratch"),
("-ab~1~", "its own segment", "abnormal, risk raised", "stop; the identity is now known"),
]
for segment, where, reading, action in SEGMENTS:
print("{:<9} {:<19} {:<30} {}".format(segment, where, reading, action))
print()
print("the parse rule is identical for all three shapes: split on the tilde, split")
print("each segment on its trailing digits, and take the last segment that is not a")
print("hash as the status. What changes is the vocabulary, which is why this file")
print("carries the segment table as data instead of hard-coding it in the parser.")
print()
print("byte-for-byte reproduction is the requirement because the edge recomputes a")
print("record hash over the body it received. Reordering two fields, shortening one")
print("length prefix, or sending an unchanged value twice are all detectable, and")
print("none of them change the structured payload you believe you are sending.")
synthetic sensor payload: 42 fields, 981 bytes in the body
round trip through the codec: byte identical, 0 parse problems
endpoint shape: POST /akam/<build>/<sensor id>/sensor_data, content-type text/plain
field category bytes origin value
--------------------------------------------------------------------------------------------------------
aj_indx index 10 counter 1491928374
aj_typ index 7 counter UNKNOWN
aj_indx_2 index 1 counter 0
session_id session 20 cookie and storage 7f3c9a11e2b84d05a6c7
session_key session 16 storage probe 4b1f0d9c8a7e6f5d
ab_score verdict 4 previous sensor verdict 0.42
platform platform 5 navigator.userAgentData Win32
platformVersion platform 6 navigator.userAgentData 15.0.0
hardware platform 2 navigator.hardwareConcurrency 12
cpu platform 17 navigator AMD Ryzen 7 5700G
browser platform 6 userAgent parse Chrome
browser_v platform 9 userAgent parse 130.0.0.0
moz browser 7 canvas 2d hash 1743902
dnt browser 11 navigator.doNotTrack unspecified
fp browser 9 canvas and webgl hash 9481273.5
websec_app_id browser 12 document.referrer host shop.example
websec_bc browser 1 cookie write test 1
websec_frame browser 1 frame ancestry 0
tz platform 15 Intl.timeZone America/Chicago
tz_off platform 5 Date offset 18000
lang platform 5 navigator.language en-US
screen platform 12 screen dimensions 1920x1080x24
sr platform 9 screen dimensions 1920x1080
ds platform 1 device pixel ratio 1
rqst timing 14 performance timeline 18|41|60|63|71
rqdd timing 14 request start deltas 12|18|21|24|26
rqurl timing 15 referrer and path /checkout/step2
m_bdr behaviour 1 pointer sample count 0
m_ti behaviour 1 pointer sample ticks 0
m_ul behaviour 1 longest pointer delta 0
m_dm behaviour 5 key, wheel and touch 0|0|0
rqet timing 13 request end timestamp 1767225600123
bcv browser 5 canvas vector crosscheck 3|0|1
wd browser 22 window and doc sizes 1920|947|1920|0|0|24|1
cloc browser 6 geolocation result denied
ckey browser 8 document.cookie probe 8c1f0a2d
fbe browser 9 storage enumeration 1|2|0|1|0
r_dat telemetry 1 response shaping check 0
r_num telemetry 1 numeric sanity probe 1
r_pos telemetry 1 position fields present 1
sw platform 6 software inventory 12|2|1
ch browser 13 UA client hint model UA-CH present
category fields bytes share what a defender pays for it
----------------------------------------------------------------------------------------------------
browser 12 104 31.8% canvas and webgl readback: forces a GPU round trip
platform 13 98 30.0% navigator and Intl reads, no extra round trip
timing 4 56 17.1% performance timeline reads, free but they leak interaction shape
session 2 36 11.0% cookie and storage reads, writes one cookie
index 3 18 5.5% loopback counters, free to compute
behaviour 4 8 2.4% event listeners, the most expensive row to store and to review
verdict 1 4 1.2% one field, but it is the one the next request is scored on
telemetry 3 3 0.9% shaping probes that exist to detect a fake navigator
origin concentration: 37 distinct browser surfaces feed these fields
counter 3 field(s)
navigator.userAgentData 2 field(s)
screen dimensions 2 field(s)
userAgent parse 2 field(s)
Date offset 1 field(s)
Intl.timeZone 1 field(s)
version fragility: what the deployed script for each build is known to emit
build wanted sent missing fields this capture does not carry
----------------------------------------------------------------------------------------------------
build 3.x 10 10 0 none, complete for this build
build 4.x 15 13 2 fp_2, audio_hash
the 4.x list wants fields this capture does not have. A payload missing them
is not rejected on syntax; it is scored on absence, and the score is what the
_abck cookie records. The codec is the cheap half. The vocabulary is the half
that has to be re-derived every time the vendor rotates the script.
_abck lifecycle: the shapes one cookie value takes across a session
first document ~hw
no payload posted yet: the pre-instrument marker
first sensor post ~-hw~m~4~3b1c9f0a2d~mr~8f21e7d4c0~
hardware segment present, status segment present
flagged post ~-hw~m~4~c0ffee11dead~mr~000000000000~-ab~1~
abnormal marker: scored, and scored badly
next document ~-hw~m~4~3b1c9f0a2d~mr~8f21e7d4c0~-ou~ok~0~mr~8f21e7d4c0~
punctuated status form: ok, plus an outbound segment the response refreshed
stale replay ~-hw~m~4~c0ffee11dead~mr~000000000000~-ou~ir~0~mr~000000000000~
re-presented unchanged: the edge leaves it flagged rather than clearing it
stage 1 is one token. Stage 2 is three segments. Stage 3 is a delimited
outbound segment that the response rewrites on every document. The lifecycle
is: no payload, then a clean status, then a punctuated refresh per request.
segment interpretation, as read from captures. The vocabulary is not
documented by the vendor and it does change between builds.
segment where it sits reading what to do
--------------------------------------------------------------------------------------------------------
~hw whole value pre-instrument marker a sensor_data POST should follow
-hw- delimited hardware segment only valid after a posted payload
~m~n~ inside -hw- mode and record length recheck the posted body size
~mr~ end of -hw- response hash changes per request; a repeat is odd
-ou- delimited outbound segment refreshed by the response, not by you
~ok~ inside -ou- status: accepted session clean, carry on
~ir~ inside -ou- status: invalid record payload rejected, re-post from scratch
-ab~1~ its own segment abnormal, risk raised stop; the identity is now known
the parse rule is identical for all three shapes: split on the tilde, split
each segment on its trailing digits, and take the last segment that is not a
hash as the status. What changes is the vocabulary, which is why this file
carries the segment table as data instead of hard-coding it in the parser.
byte-for-byte reproduction is the requirement because the edge recomputes a
record hash over the body it received. Reordering two fields, shortening one
length prefix, or sending an unchanged value twice are all detectable, and
none of them change the structured payload you believe you are sending.
Reading the Segments Without Guessing
The segment table gives the parse rule: split on the tilde, split each segment on its trailing digits, and take the last segment that is not a hash as the status. That rule is stable across builds. The vocabulary around it is not.
~ok~ inside the outbound segment means the session is clean. ~ir~ means the
payload was rejected and a fresh post is required. The abnormal segment is the
one to treat as terminal. The response hash sits at the end of the hardware
segment and changes per request, so a repeated value in a capture is worth
looking at: it means a request was served from a cache or a replayed body.
Because the vocabulary changes, the snippet carries it as data rather than hard-coding it in the parser. A parser that special-cases segment names has to be edited on every rotation; a parser driven by a table has to be edited only when a segment appears that the table does not describe. The same principle applies to the field list: keep the vocabulary in a data file you can diff, and keep the parser generic.
Field Inventories Are Worth Keeping as Artefacts
The inventory output is longer than it looks. It is a diffable record of what your environment reports, in bytes, per field, with the origin for each. When a session starts getting flagged after a browser upgrade, the inventory is what tells you which field changed, and it takes seconds to compare two of them.
The same artefact argues for storing behaviour and platform fields separately in your own telemetry. The category split tells you which of your fields are worth transporting on every request and which are worth sampling, and the byte column tells you what that decision costs. Keeping the inventory under version control also turns an upgrade from a debugging session into a diff review, because the change is visible before the score moves rather than after.
What Defenders Reason About, and What That Implies
From the defender's side the payload is a bundle of consistency claims: the
platform fields must agree with each other, the timing deltas must be possible for
a human-speed interaction, the session identifiers must match the jar, and the
shaping probes must return what a real implementation returns. The probes at the
bottom of the inventory exist for one reason: to detect a navigator that was
patched rather than a browser that is genuinely old.
The implication for a client is narrow. Consistency across fields is the part you cannot fake piecemeal; a claim set that contradicts itself is cheaper to detect than any single false claim, and defenders are optimising for exactly that. A single wrong field is noise. A field that disagrees with three others is a verdict, and the disagreement is usually the thing that survives every attempt to patch the individual field.
Checklist
- Treat the body as a byte sequence with a length prefix, and verify a round trip before anything else.
- Keep the field inventory with bytes and origins, and diff it across upgrades.
- Record the build identifier the payload was built for; the vocabulary moves.
- Read the cookie as the previous verdict, and read the body as the evidence.
- Treat the abnormal segment as terminal for that identity; start a new one.
- Compare response hashes across requests; a repeat means a replay or a cache.
- Expect scored-on-absence: a missing field is a finding, not a neutral, and it will not raise an error you can catch.
- Keep the segment and field vocabularies in data files so a rotation is a diff rather than a code change.
The Legitimate Route
Analysing this on a site you own is normal security work, and the vendor's own documentation and test keys cover the supported integrations. For collecting data you have a right to, an official or partner feed removes the sensor system from the picture entirely, and where a sensor stands in the way of an authorised test, a written scope that names the endpoints and the rate is a far better instrument than a forged payload. Rebuilding a payload for a site you do not own, to get at content you are not entitled to, is the line this lesson does not cross; the analysis is worth doing on a target where you already have permission. A useful habit is to keep the synthetic payload and the field inventory in the same repository as your own test target, so the distance between studying the structure and weaponising it against a third party stays visible.