About
How this was built, and where it is weak
Stated plainly, because a market map you cannot audit is worth very little.
Collection. Public job-board APIs — Greenhouse, Ashby and Lever — queried directly in August 2026. 384 company boards probed, 189 resolved, 19,755 live postings retrieved across five sweeps: general, infrastructure and developer tooling, security, identity, and IT infrastructure. No scraping, no logged-in sources, nothing behind a terms-of-service prohibition.
Selection. Postings matched by title against per-family keyword patterns: 1,558 AI · 1,405 platform · 1,060 security · 182 identity · 670 IT infrastructure. Full description text parsed for 1,800 of them, and term frequencies computed by regular expression over that text.
Compensation. Extracted from description body text by identical method for every family. US base salary only. Per-archetype bands are an apportionment by title and seniority, not directly measured per archetype.
Interview loops come from published 2026 guides and candidate reports, not from the posting corpus, and are the least certain content here.
Reproducibility. Harvest scripts are in tools/; the raw August 2026
snapshot is committed under data/raw/, gzipped. Every figure can be recomputed. Re-running
the harvest produces a different snapshot, because the market moves.
Five limits that matter
- The India sample is small and unrepresentative. Naukri, Workday-hosted employers and own-portal companies are absent, excluding most large GCCs and most of the services sector. Treat every India number as directional only.
- Compensation is biased upward — bands are published mostly where US state law requires it, so the median is a median of the transparent.
- Title matching is imperfect both ways. Terms like “model”, “agent” and “identity” pull false positives; real work under a plain “Software Engineer” title is missed.
- Term frequency measures mention, not importance.
- The “AI mentioned” figures use a deliberately loose pattern. The 92% for security counts any mention including boilerplate; the strict measure — postings naming actual AI risk practice — is 9%. Both are reported because the gap is the finding.
An independent control
Every figure here comes from one collection method, which is a weakness. A reader supplied a second sample — 48 postings gathered by hand from LinkedIn, weighted toward enterprise and financial services — and the identical analysis was run over it.
The central claim survived. PyTorch, PhD, distributed training and publications stayed in single digits in both corpora while agents and customer-facing work dominated both. One figure was corrected: retrieval appears in 11% of AI-native postings and 72% of enterprise ones, and this site had understated it. One was missing entirely: responsible AI and governance language appears in 87% of the enterprise sample and had not been measured at all.
Both the confirmation and the correction are written up in Which market are you interviewing for? — including the caveat that 48 hand-collected postings are directional rather than representative.
Why not LinkedIn or Naukri
Automated collection is what both prohibit — a reader manually saving postings they viewed is a different thing, and that is where the control sample above came from. Indian law is also materially less scraper-friendly than US law — the IT Act 2000 §43(b) creates civil liability for copying data from a computer resource without permission, with no requirement to show a technical barrier was defeated.
It is also unnecessary. Both are mirrors: postings originate in Greenhouse, Ashby, Lever and Workday, which serve structured data publicly to their own front ends.