Lakshya
Lakshya Engineering roles — India & United States Compiled 21 August 2026

What hiring
actually asks for.

Fourteen thousand live requisitions, read end to end, across two role families. Both produced the same shape of surprise: the loudest part of each field is not the part that is hiring.

14,011Postings read
122Company boards
2,963Roles matched
1,091Full JDs parsed
14Role archetypes

The finding

The market bifurcated. Prep didn't follow.

Parsing the full text of 592 US job descriptions produces a ranking that is hard to reconcile with how people study for these jobs. The three most frequently named things are not modelling skills at all. They are agents and tool use, evaluation, and working directly with customers.

Applied track — the majority
64%name agents, tool use, or function calling
59%name evaluation or benchmarking
58%are explicitly customer-facing
30%name ambiguity as a job requirement
Frontier track — the minority
11%name PyTorch
11%ask for a PhD
6%name distributed training
4%name RLHF, and 4% name CUDA

Read those columns against each other. A candidate grinding PyTorch, distributed training and paper reproductions is preparing for the column that appears in roughly one JD in ten. Meanwhile the skill named in 59% of postings — evaluation — has a dedicated job title in fewer than 2% of them, which means almost nobody is studying it and almost every hiring manager is screening for it.

The single most useful correction

If you have no research track record, the frontier track is a poor use of preparation time — not because it is beneath you, but because it is a small, credential-gated slice of the market that you cannot enter by studying harder for six weeks. The applied track is large, growing, pays comparably at senior levels, and rewards exactly the thing a working engineer already has: shipped systems and the scars from them.

One more structural fact worth internalising before you read further. Of 1,558 AI roles, 313 were Staff-level and 235 Senior — against one new-grad posting and 18 internships. This is not an entry-level market. It is a market buying judgement, and every interview loop below is built to test judgement rather than recall.

Evidence

What the job descriptions actually name

Share of job descriptions whose full text mentions each term. Limited to the terms scanned identically in both markets, so the two bars are directly comparable.

United States — n=592 India — n=72
agents / tool use
64%
54%
evals / benchmarks
59%
45%
customer-facing
58%
44%
Python
48%
68%
MLOps / registry
18%
18%
fine-tuning
13%
25%
RAG / retrieval
11%
31%
PyTorch
11%
23%
PhD
11%
16%
SQL
11%
15%
distributed training
6%
9%
publications
6%
4%

Terms scanned in the US corpus only, for completeness: mentorship 38%, observability 33%, AWS 30%, GCP 27%, Azure 25%, TypeScript or React 22%, Spark 22%, Go 14%, Kubernetes 14%, embeddings or vector databases 11%, Model Context Protocol 8%, prompt engineering 7%, Rust 6%, inference optimisation 5%, greenfield or zero-to-one 5%.

Two of those deserve a second look. Kubernetes at 14% and Terraform at 4% tell you AI roles are not infrastructure roles in disguise — the platform work is real but concentrated in one archetype. And MCP at 8% is the fastest-moving term in the corpus; a year ago it was absent. Naming it fluently is currently a cheap differentiator.

The seven archetypes

Seven jobs wearing the same word

Every archetype below is described in the same nine parts, so you can read one closely and skim the rest. The coloured chip tells you which track it belongs to — that is the single most important thing to know before you spend a month preparing.

Applied track Largest single title family

1 · Forward Deployed Engineer / Applied AI Engineer

The biggest cluster in the entire dataset, and the one with the least prep material written for it.

What the job actually is

You are embedded with a customer, and you write the code that makes a general model work inside their specific mess — their schemas, their auth, their compliance boundary, their idea of what "correct" means. Roughly half the work is engineering and half is consulting: figuring out what the customer actually needs versus what they asked for. The role exists because frontier models are general and enterprises are not, and that gap does not close by writing better documentation.

Titles this hides behind

Forward Deployed EngineerApplied AI Engineer Applied AI ArchitectAI Deployment Engineer AI Engineer — FDEPartner AI Deployment Engineer Forward Deployed AI EngineerSolutions Engineer (AI)

Who is hiring it

OpenAI, Anthropic, Palantir, Sierra, Harvey, Glean, Scale AI, Writer, Instabase, ElevenLabs. Frequently split by vertical — financial services, public sector, healthcare and life sciences, retail, manufacturing, media and games — which is a useful signal: pick the vertical where you already speak the language. In India the same shape appears at Sarvam, Turing and Observe.AI.

How to recognise it in a JD

Look for customer-facing, onsite, travel, stakeholder, engagement, ambiguity, and prototype to production. If the requirements list names a vertical industry before it names a language, you are reading an FDE posting regardless of what the title says.

The loop

Recruiter screen Motivation and travel tolerance. They are checking you understand this is not a pure IC role.
Hiring manager conversation Your history of owning outcomes, not tickets. Have a story where you changed what was being built.
Technical deep dive Real production LLM systems — RAG, evals, agents, fine-tuning trade-offs. Depth in one area is enough.
Ambiguous customer case study — the decisive round 45–60 minutes. A vague customer problem, and you decompose it into a plan out loud. Reported as the lowest pass rate in the loop, at roughly 40%, and the heaviest weighted at around 30%.
Behavioural / values Conflict with a customer, a project you killed, a deadline you missed and what you said.

Five to eight stages across three to six weeks. At OpenAI and Palantir, roughly half the total evaluation weight sits on non-coding rounds.

Prepare like this

  • Build one agent against genuinely messy data — not a clean demo dataset. Inconsistent schemas, missing fields, three date formats. The stories that land in the case study round come from having actually hit this.
  • Practise scoping out loud, on a timer. Take a one-line business problem, spend 45 minutes talking through decomposition to a recording. Listen back. Most candidates fail this round by solving too early rather than by solving wrong.
  • Prepare three "they asked for X, the real problem was Y" stories with the discovery moment in each, and what it cost to have found it late.
  • Be able to whiteboard an enterprise integration: where the data lives, how auth flows, what happens to PII, what the latency budget is, who pays for tokens, and what runs inside their VPC.
  • Know the failure modes cold — retrieval returning plausible-wrong context, tool calls failing silently, cost blowing up on retries.

Be ready for

  • "A customer says the model is wrong 30% of the time. What do you do first?" They want you to ask what "wrong" means and how it was measured, before proposing anything.
  • "You have two weeks and an executive demo. What do you cut?" Testing whether you scope to a demonstrable slice or promise everything.
  • "The customer's data is far worse than they told us. Walk me through the conversation." Testing whether you can deliver bad news without either sugar-coating or blaming.
  • "How would you know this deployment succeeded, six months from now?" Testing whether you think in outcomes or in shipped features.

Red flags in the posting

Pre-sales in disguise. If the JD never mentions a codebase, code review, or shipping, this is a solutions consultant with an engineering title and an engineering interview but not an engineering career path.

Unbounded travel. "Travel as needed" with no percentage. Ask for the number in the recruiter screen; a real answer is 20–40%, and evasion is informative.

No engineering manager in the loop. If every interviewer is from sales or customer success, your performance reviews will be too.

Compensation

United States — base$180k – $310k Observed range for the archetype within a corpus median of $240k. Equity is the larger lever at frontier labs.
India — total₹35L – ₹70L Estimated. No Indian JD in the corpus published a number — see method note.
Applied track Fastest-growing

2 · Agent / AI Product Engineer

Named in 64% of all job descriptions. The centre of gravity of the whole market.

What the job actually is

You build product features where a model takes actions — calling tools, holding state across turns, recovering from its own mistakes. The hard part is not the prompt. It is everything around it: what happens when the third tool call in a chain returns garbage, how you keep a multi-turn session inside a context budget, how you stop a retry loop from costing $400, and how you prove a change made things better rather than merely different.

Titles this hides behind

Software Engineer, AgentsAI Agent Engineer Frontier Agents EngineerAI Engineer Software Engineer, AI PlatformSDE III — Gen AI Software Engineer, AI ProductAgent Builder

Who is hiring it

Effectively everyone — this is the archetype that spread furthest beyond AI-native companies. Sierra, Glean, Writer, Cursor, Notion, Linear, Vanta, Ramp, Postman, Atlan, plus the whole frontier-lab product surface. In India: Postman, Sarvam, Observe.AI, Tekion, Meesho.

How to recognise it in a JD

Look for agentic, tool use, function calling, orchestration, multi-turn, guardrails, and increasingly MCP. The tell that separates it from archetype 3 is that the metrics named are product metrics — task completion, resolution rate, deflection — rather than model metrics.

The loop

Recruiter screen What you have actually shipped that a user touched.
Practical coding Frequently "build a small agent" rather than an algorithms puzzle. An IDE, real API docs, 60–90 minutes.
LLM system design — the decisive round RAG fundamentals are the entry fee, not the bar. The real question starts one level up: what happens when the system you built breaks in a way no tutorial covers.
Product sense Where a model should and should not be in the loop, and what quality bar ships.
Behavioural At senior and staff level, interviewers pick three to five topics and drill into failure modes and what went wrong last time.

Prepare like this

  • Ship one agent with real tool calling and a real eval harness. Not a notebook — something with retries, timeouts, structured logs, and a regression suite you can run before and after a prompt change.
  • Break it deliberately and write down what happened. Kill a tool mid-chain. Return malformed JSON. Overflow the context. The interview is largely about failure modes, and you cannot narrate failures you have not seen.
  • Learn MCP properly while it is still a differentiator at 8% JD penetration and climbing.
  • Get fluent in cost and latency arithmetic — tokens per request, cost per resolved task, p95 versus p50 when a chain has four sequential calls.
  • Have a position on prompt injection and what you actually do about it when the agent has write access to something.

Be ready for

  • "Your agent calls a tool that fails silently and returns an empty result. How do you detect it?" The best answers reach for evals and tracing, not for a better prompt.
  • "How do you evaluate a multi-turn agent, where there is no single correct output?" The archetype-defining question. See archetype 6.
  • "The context window fills up mid-task. Now what?" Summarisation, retrieval, handoff, or task decomposition — and the cost of each.
  • "You changed a prompt and the demo looks better. Do you ship it?" Correct answer involves a golden set and a regression run, not a vibe.

Red flags in the posting

No mention of evaluation anywhere. A team building agents without evals is shipping on vibes and will blame you when quality regresses invisibly.

A framework named as the job. "Experience with LangChain required" as a headline requirement usually means the team has confused a library with an architecture.

Compensation

United States — base$170k – $300k Wide, because the title spans product teams at $170k and frontier labs well past $300k.
India — total₹30L – ₹65L Estimated. A 20–40% premium over generalist ML roles is widely reported for hands-on LLM work.
Applied track Strongest in India

3 · ML Engineer — Ranking, RecSys, Risk

The classical role. Still large, especially in India, and the one most often rebranded rather than changed.

What the job actually is

Models that serve live traffic: what to show, what to rank, what to block. Feed ordering, search relevance, ads, fraud and credit decisioning. The LLM content in these roles is frequently a paragraph bolted onto a job that is otherwise unchanged since 2022 — which is not a criticism, because this is where a great deal of measurable business value still sits, but it does change how you should prepare.

Titles this hides behind

Machine Learning EngineerStaff MLE Applied ScientistApplied Scientist — Recommendations Data Scientist IIISDE II — AI Senior MLE — NLPML Engineer, Vision

Note the "Applied Scientist" convention — an Amazon lineage, and in India it very often means ranking and recommendations specifically rather than research.

Who is hiring it

India is the stronghold: Meesho, InMobi, Glance, PhonePe, Flipkart, Swiggy, Navi, Slice. In the US: Reddit, Pinterest, DoorDash, Instacart, Robinhood, Coinbase. Recommendations and ranking appear in 38% of Indian AI JDs against a far smaller share of US ones — the clearest geographic difference in the entire dataset.

How to recognise it in a JD

Look for ranking, personalisation, relevance, CTR, A/B, feature store, Spark, underwriting. If Spark and SQL appear near the top, this is archetype 3 whatever the title promises.

The loop

Coding screen Still genuinely algorithmic here, unlike archetypes 1 and 2. Do not skip this preparation.
ML depth Features, leakage, class imbalance, calibration, why your metric moved.
ML system design — the decisive round Design a ranking or fraud system end to end: candidate generation, features, training cadence, serving, and the offline-to-online gap.
Experimentation and metrics A/B design, power, novelty effects, and what you do when offline gains do not appear online.
Behavioural

Prepare like this

  • Own one end-to-end story with real numbers — baseline, change, measured lift, and how you knew it was the change rather than seasonality. Vague numbers read as borrowed credit.
  • Be sharp on the offline-online gap, because it is the round most candidates lose. Training-serving skew, feedback loops, position bias in logged data.
  • Separate model metrics from business metrics out loud. NDCG went up and revenue did not — be able to explain how that happens and what you do next.
  • Refresh the classical material. This is the one archetype where gradient boosting, calibration and feature engineering still carry an interview.

Be ready for

  • "Offline AUC improved, the A/B was flat. What are your hypotheses, in order?" The signature question of this archetype.
  • "Your training data comes from what your current model showed users. What is the problem?" Position and exposure bias — and whether you have thought about counterfactual logging.
  • "How often do you retrain, and how would you know that is right?"

Red flags in the posting

An LLM paragraph with no LLM substance. If "GenAI" appears once in the responsibilities and never again in the requirements, expect the actual work to be gradient boosting on tabular data. Fine if you want that — costly if you took the role to move into LLM work.

No mention of experimentation. A ranking team without A/B infrastructure cannot tell you whether your work mattered, and neither can your promotion packet.

Compensation

United States — base$165k – $260k Sits below archetypes 1, 2 and 5 at equivalent level in the observed corpus.
India — total₹15L – ₹40L mid · ₹30L – ₹60L senior Bengaluru senior AI/ML reported around ₹22L average, with a ₹13L–₹32L interquartile range.
Frontier track ≈10% of the market

4 · Research Engineer / Scientist / Member of Technical Staff

The role everyone pictures. A tenth of the postings, and the hardest to enter laterally.

What the job actually is

Making empirical progress on model capability under a compute budget. In practice that is far more engineering than the word "research" suggests — data pipelines, training runs that fail at hour nine, evaluation harnesses, and a great deal of careful measurement. Member of Technical Staff is worth understanding as a convention rather than a level: several labs use it deliberately flat, so the title carries no seniority signal and you must ask about scope directly.

Titles this hides behind

Member of Technical StaffResearch Engineer Research ScientistML Researcher, Foundational Models Staff Research Engineer — CodeAI Research Engineer

Who is hiring it

Anthropic, OpenAI, Cohere, Mistral, and — the notable Indian entry — Sarvam, which is posting genuine foundation-model roles: ML Researcher for foundational models, ML Engineer for training infrastructure, and ML Engineer for data. That is new. Until recently, frontier research roles were essentially not available in India.

How to recognise it in a JD

Look for publications, NeurIPS, ICML, pre-training, RLHF, ablation, scaling laws. Only 6% of the corpus mentions publications at all, so when it appears it is a real filter rather than boilerplate.

The loop

Recruiter screen
Research coding screen Frequently a training loop or a data transform written from scratch, not a puzzle.
Deep dive on your own work — the decisive round Ninety minutes on one project you did. Expect three levels of "why did you choose that" on every decision, including ones you made casually.
Pair programming on a research task Open-ended, with the interviewer watching how you form and discard hypotheses.
Values / alignment discussion At safety-focused labs this is substantive and can be disqualifying, not a formality.

Prepare like this

  • Reproduce one paper end to end and be able to defend every hyperparameter. A single deeply-owned reproduction beats five shallow ones.
  • Interrogate your own past work to three levels before the deep dive. Why that architecture, why that baseline, why you believed the result.
  • Write a training loop from scratch in PyTorch or JAX, no trainer abstraction, until it is muscle memory.
  • Be honest about negative results. "We tried it, it did not work, here is what we learned" is a strong answer here and a weak one almost nowhere else.

Be ready for

  • "You have 100 GPU-hours and a hypothesis. Design the experiment." Testing whether you can scope research to a budget — the actual daily constraint.
  • "Your eval improved by 2 points. Convince me that is real." Variance, seeds, contamination, and whether the eval measures what you claim.
  • "What is the strongest argument against your own last result?"

Read this before you commit a month to it

This is the highest-effort, lowest-yield archetype for a lateral candidate. It is roughly a tenth of the market, screens partly on credentials you either have or do not, and competes against people with full-time research track records. If you have publications or a research-adjacent history, pursue it. If you do not, the same month invested in archetypes 1, 2 or 6 has a materially better expected outcome — and those roles pay comparably.

Compensation

United States — base$240k – $400k+ The top of the published range, and equity dominates total compensation at frontier labs.
India — total₹40L – ₹90L Estimated, thin sample. Sarvam is close to the only domestic data point.
Frontier track Small · most durable

5 · AI Platform & Inference Infrastructure Engineer

Where the GPUs, the throughput and the cost-per-token live. Fewer postings, less competition per posting.

What the job actually is

Making training and inference fast and affordable at scale. Cluster scheduling, distributed training that survives node failure, serving stacks, batching strategy, quantisation, and the unglamorous arithmetic of how many tokens per second you get per dollar. This is the archetype whose skills age best — the specific models change yearly, the systems constraints do not.

Titles this hides behind

Platform Engineer — AI Infrastructure Software Engineer, AI PlatformPerformance Engineer, Inference ML Engineer (Training Infra)Performance Engineer, On-Device Inference ML Ops Engineer

Who is hiring it

Baseten, Modal, Together, Fireworks, Databricks, Snowflake, plus the infrastructure organisation inside every frontier lab. In India, Sarvam is hiring both training infrastructure and inference performance, and Glance is hiring on-device inference — which is a distinctly Indian specialism driven by a mobile-first market.

How to recognise it in a JD

Look for throughput, latency, quantisation, KV cache, vLLM, Triton, CUDA, FSDP, multi-node, cost per token. These terms are rare corpus-wide — CUDA 4%, distributed training 6% — which is precisely why they are a clean signal when present.

The loop

Systems coding screen Concurrency, memory, profiling. Closer to a classical backend screen than an ML one.
Distributed systems design — the decisive round Design a serving stack to a latency and cost target. Expect to be pushed on the arithmetic until it either holds or does not.
Performance deep dive A time you made something faster. They will ask how you measured, before and after.
Behavioural / on-call reality On-call appears in only 5% of AI JDs corpus-wide, but disproportionately in this archetype. Ask about rotation.

Prepare like this

  • Do the memory arithmetic until it is instant — parameters, optimiser state, activations, and KV cache growth as a function of batch and sequence length. This gets asked directly and answered badly.
  • Know continuous batching, paged attention, speculative decoding, and tensor versus pipeline parallelism well enough to explain the trade-off, not just the name.
  • Bring one measured win. A latency or cost-per-token improvement with a before, an after, and the profiling that found it.
  • Serve a real model yourself on rented GPUs and hit an actual throughput target. The gap between reading and doing shows immediately here.

Be ready for

  • "Estimate the GPU memory to serve a 70B model at batch 32, 8k context." Arithmetic, out loud, with your assumptions stated.
  • "p50 latency is fine, p99 is terrible. Where do you look?" Queueing, batching, preemption, and cold starts.
  • "Cut inference cost 40% without hurting quality. What do you try, in what order?"

Red flags in the posting

"AI Platform" that is actually Kubernetes. Some postings under this title are general platform engineering with no model-specific work at all. Ask in the screen whether the team owns the serving stack or merely the cluster it runs on.

Compensation

United States — base$200k – $350k Scarcity-priced. The corpus 90th percentile of $350k is populated heavily by this archetype and archetype 4.
India — total₹35L – ₹75L Estimated. Very few domestic postings, so the range is wide and weakly supported.
Applied track Named 59% · titled <2%

6 · AI Evaluation & Quality

The largest gap between what is demanded and what is studied. Read this section even if you are targeting another archetype.

What the job actually is

Deciding whether a non-deterministic system got better or worse, with enough rigour that a team can ship on the answer. Golden sets, regression suites, human labelling protocols, LLM-as-judge and its many biases, and the statistics of drawing conclusions from a few hundred examples. It is measurement science applied to a system that gives a different answer every time you ask.

The arbitrage is stark: 59% of job descriptions name evaluation, and under 2% have it as a job title. Almost nobody prepares for it, and it is assessed in every applied-track loop. Getting good at this raises your performance in archetypes 1, 2 and 3 simultaneously — which makes it the highest leverage thing on this page.

Titles this hides behind

Data Scientist — EvaluationsAI Quality Engineer SDET — AI EvaluationApplied Scientist, Evaluation Research Engineer, Evals

Prepare like this

  • Build a real eval harness for something you have shipped. A golden set, a scoring function, a regression run, a report. Fifty well-chosen examples beat five thousand scraped ones.
  • Learn the LLM-as-judge failure modes by name: position bias, verbosity bias, self-preference, and score compression at the top of the scale. Then learn the mitigations — swapping order, pairwise comparison, calibrating against human labels.
  • Always carry a control. A number without a baseline is not a result. If your agent resolves 71% of tickets, the interview question is 71% against what — the previous version, a human, or a trivial keyword rule that would have got 65%. This single habit distinguishes strong candidates more reliably than any technical depth.
  • Know your statistical power. On 200 examples, a 3-point difference is usually noise. Be able to say so, with the arithmetic.
  • Design a human labelling protocol including the disagreement rate you would accept and what you do when annotators disagree.

Be ready for

  • "How do you evaluate something with no single correct answer?" Rubrics, pairwise preference, and task-completion proxies — plus honesty about what each misses.
  • "Your LLM judge agrees with humans 85% of the time. Is that good enough to ship on?" Depends entirely on where the 15% falls. Strong answers ask about the error distribution.
  • "How do you stop your golden set from going stale?"
  • "You improved the eval score and users complained more. Explain." The eval measured the wrong thing — and whether you can say that plainly is the test.

Compensation

United States — base$160k – $260k Underpriced relative to demand today, which is exactly what an emerging title looks like.
India — total₹25L – ₹55L Estimated. Sarvam and Postman both post evaluation-specific roles.
Commercial track Newest titles

7 · AI Product Manager / GTM Engineer / Engagement Lead

The commercial surface of AI, and it is becoming genuinely technical.

What the job actually is

Owning an AI capability as a business outcome rather than a feature — what ships, at what quality bar, priced how. The interesting development is GTM Engineer, a title that barely existed two years ago and now appears across the corpus: someone who writes code in service of revenue, building the demos, integrations and internal tooling that close deals. It sits between archetypes 1 and 7 and is a genuinely good landing spot for an engineer who likes customers.

Titles this hides behind

Product Manager, Agent DevelopmentAI Product Manager GTM EngineerAI Engagement Lead Engagement Manager — AI AgentsStrategist, Agent Development Director, AI TransformationAI Adoption Lead

The loop

Recruiter screen
Product sense Scoping a product whose core component is non-deterministic. This is the part that differs from ordinary PM interviews.
Technical fluency Not coding, but you must hold your own on context windows, latency, cost and why the model is confidently wrong.
Strategy case — the decisive round Build versus buy, pricing under variable inference cost, and what quality bar justifies launch.
Executive conversation

Prepare like this

  • Have a position on eval-gated launch — the specific bar you would hold a release to, and what you would do the week it is missed.
  • Know token unit economics well enough to price a feature. Cost per request, per resolved task, and how it moves when the agent retries.
  • Be able to say what you would cut. Every strategy case is secretly a scoping test.
  • Prepare the failure story: an AI feature that shipped and underperformed, and what you learned about the gap between demo and production.

Red flags in the posting

"AI Transformation" with no product surface. Several postings under this banner are internal change management with no artefact to point at afterwards. Ask what shipped last quarter and who used it.

Compensation

United States — base$170k – $290k GTM Engineer frequently carries variable compensation — clarify the split early.
India — total₹30L – ₹70L Estimated, and unusually wide because the title spans PM and quota-carrying roles.

Geography

India and the US are not hiring the same job

DimensionUnited StatesIndiaWhat it means for you
Dominant archetypeForward Deployed & Agent Engineering ML Engineering & Applied Science Indian loops still weight classical ML depth more heavily.
RecSys / rankingMinor38% of JDs Consumer scale is the Indian speciality. Lead with ranking experience.
PyTorch named11%23% Indian roles build models; US roles increasingly orchestrate them.
RAG named11%31% RAG is still a differentiator in India and already table stakes in the US.
Customer-facing58%44% High in both. The "AI job means no meetings" assumption is wrong in both markets.
Frontier researchEstablishedEmerging — Sarvam Domestic foundation-model roles now exist. They did not two years ago.
On-device / edge AINicheA distinct theme Glance and Sarvam both hire for it. A mobile-first market creates a real specialism.
Published compensation1,237 figures0 figures You must anchor externally and raise money early in Indian processes.

The compensation asymmetry is actionable

Across 592 US job descriptions the corpus contained 1,237 published salary figures — 10th percentile $166k, median $240k, 90th percentile $350k in base pay. Across 72 Indian job descriptions it contained zero. That is not a data collection artefact; Indian postings simply do not publish bands.

The practical consequence: in a US process the band is public and negotiation is about where you land within it. In an Indian process the band is private, the first number is frequently anchored to your current salary, and the single highest-leverage move is to decline to disclose current compensation and state a target instead. External reference points for 2026 put entry at ₹6–12L, mid at ₹15–30L and senior at ₹30–60L, with Bengaluru senior AI/ML averaging around ₹22L and a reported 20–40% premium for demonstrable hands-on LLM, RAG or LLMOps work.

Preparation

A three-week protocol

Deliberately archetype-agnostic. Substitute the specifics from your target section above. It assumes roughly ten hours a week, which is what an employed person actually has.

Week 1 — Build the artefact

One real, working thing you can screen-share. An agent, a ranking pipeline, a served model, an eval harness — matched to your archetype. It must touch messy data and it must be broken and fixed at least once.

The artefact is not portfolio decoration. It is the source of every specific answer you will give in weeks 2 and 3.

Week 2 — Instrument and measure

Add evaluation to whatever you built. Golden set, baseline, a control, and a number you can defend. Deliberately break it three ways and write down what each failure looked like from the outside.

This week is what separates candidates. Most people stop after week 1 and then cannot answer "how did you know it worked".

Week 3 — Rehearse out loud

Record yourself doing the decisive round for your archetype — the case study, the system design, the deep dive — on a timer, and watch it back. Then write your five stories: an outcome you owned, a failure, a disagreement, a thing you cut, and a measurement that surprised you.

Watching your own recording is unpleasant and is the single highest-yield hour in the three weeks.

Method & limits

How this was built, and where it is weak

Stated plainly, because a market map you cannot audit is worth very little.

Collection. Public job-board APIs — Greenhouse, Ashby and Lever — queried directly on 21 August 2026. 128 company boards probed, 81 resolved, 11,229 live postings retrieved. No scraping and no logged-in sources.

Second sweep. A further 94 infrastructure and developer-tooling boards were probed for the platform family, of which 41 resolved, adding 2,782 postings for a corpus total of 14,011 across 122 boards.

Selection. 1,558 postings matched an AI keyword pattern on title, and 1,405 matched a platform, SRE, DevEx or SDLC pattern. Full description text was then parsed for 592 US and 72 India AI postings, plus 427 platform-family postings, and term frequencies computed by regular expression over that text.

Compensation. Extracted from job description body text by identical method for both families. US base salary only — n=1,237 for AI, n=515 for platform. Per-archetype bands are my apportionment of those corpus figures across archetypes using title and seniority, not directly measured per archetype; treat them as guides rather than as data.

Four limits that matter

  • The India sample is small and unrepresentative. 72 AI roles from 11 boards. Naukri, Workday-hosted employers, and companies on their own portals are entirely absent — which excludes most large GCCs, the Indian arms of Google, Microsoft and Amazon, and most of the services sector. Treat every India number as directional only. The US figures are far better supported.
  • Compensation is biased upward. Companies publish bands mostly because California, New York, Washington and Colorado require it. That over-samples high-paying coastal employers, so the $240k median is a median of the transparent, not of the market.
  • Title matching is imperfect in both directions. Terms like "model" and "agent" pull in false positives, while real AI work under a plain "Software Engineer" title is missed entirely.
  • Term frequency measures mention, not importance. A JD naming Kubernetes once is counted the same as one built around it. The percentages show what the market talks about, which is related to but not identical with what it tests.

Interview-loop structure and stage weightings are drawn from published 2026 interview guides and candidate reports, not from the posting corpus, and are the least certain content on this page. Treat stage counts as typical rather than exact, and expect variation by company and by level.

Next

Families still to map

Two families mapped, and the method transfers directly. Each remaining family needs its own archetype pass, because the titles collide within a family in exactly the way "AI Engineer" and "Platform Engineer" collide within these two.

Cyber Security EngineeringAppSec, product security, detection engineering, cloud security, and the new AI attack surface.
Security Risk & GovernanceRisk analysis, GRC, third-party risk, and the control frameworks behind them.
Internal Audit & AI AssuranceIT audit, model risk management, and auditing AI systems — a genuinely new discipline.
Identity & Access ManagementIAM engineering and architecture: directory, federation, SSO, lifecycle, entitlements.
Privileged Access ManagementCyberArk and HashiCorp Vault as a career track in its own right, independent of AI.
Data Platform & Analytics EngineeringThe family most reshaped by the shift from pipelines to contracts.
Engineering ManagementEM against Tech Lead against Staff IC, and how the loops genuinely differ.

And a companion handbook

Every archetype on this page shares a spine — system design under constraint, measuring a thing that resists measurement, incident and failure narrative, and the arithmetic of cost and latency. Those belong in one deep, diagrammed handbook rather than repeated fourteen times. That is the next build, and it is what turns this atlas from a map into preparation.