Andrew Ng · The AI Engineering Skills Map · 2026-08-15
Four core skills. Three met,
one above the bar.
Ng derived four core AI engineering skills, plus one underlying mindset, from an analysis of over 10,000 job postings, dozens of structured interviews and survey data. This report checks Miao YU against each of them. The primary evidence is machine-checkable traces, not her own account of herself.
Her strongest skill happens to be the one Ng says every developer now needs and most people still treat as mere usage: governing a coding agent as a system that will fail.
00 · Origin
This checkup started with the post below. It passed a million views on the day it went up, and most replies asked "what should I learn?". We asked the question the other way round: someone already working like this, how many of the four does she actually meet?
assets/ng-tweet.png01 · How the evidence was gathered
This checkup deliberately does not rest on what she says about herself. The principle is hers: what I feed you may be partial, it is what I chose to tell you — but the way I work with you cannot be faked.
Three tiers of evidence
Tier 1 · machine-checkable traces (the primary basis). With her authorisation I scanned every repository in her workspace and her ~/.claude configuration: commit counts, timestamps, gate scripts, eval results, hook configuration, worktree structure. None of it passes through anybody's narration.
Tier 2 · behavioural samples from our collaboration (section 06). What she actually asked for, corrected and rejected. This class of evidence cannot be staged, because at the time she did not know it would become a document.
Tier 3 · her own account. Industry experience, patents, papers and competition results from before 2026. I have no independent source for these, so they are excluded from every score. If they hold up, this profile only gets stronger.
Counting and quoting
Only commits authored by her are counted, deduplicated by commit hash across branches. Upstream commits written by anyone else are excluded. A report that counts other people's commits as personal output makes its remaining numbers worthless.
Ng is quoted in key sentences rather than reproduced at length: each dimension carries one or two sentences of his original wording with a link back to the source, so the context can be checked.
02 · Machine-checkable traces
All figures come from a scan; the method for each is given alongside, so anyone can count them again.
| Commits authored by her16 repositories, deduplicated by hash across branches | 1,751 |
| Active repositoriesrepositories containing commits of hers | 16 |
| Observation windowfirst to last commit, 2026-07-13 → 08-15 | 33 days |
| Repos with a CLAUDE.md governance filedirectory-level rules, not a global template | 7 |
Repos with CI workflows.github/workflows/ | 8 |
Repos with a deterministic gate scriptscripts/run_checks.py and similar | 6 |
| Gates in a single repomiao-yu-lab; the same command runs locally and in CI | 15 |
Architecture decision recordsdesign/0001–0011 | 11 |
Skill eval runsmurDrift/evals/results/, all on 2026-08-04 | 57 |
| Parallel worktreesacross 4 repos; names show manager/executor separation | 15 |
Self-authored user-level skills~/.claude/skills/ | 7 |
Memory-mirror mount points~/.claude/memory-mirror.map | 16 |
Thirty-three days, 1,751 commits, ten repositories moving in parallel. The speed is not the point. The point is that gates, ADRs, evals and handoff protocols grew inside those same repositories over the same period. Speed and governance usually trade against each other; here they did not.
03 · Scores
The scores are my judgement, not a certification. The basis for each one is in the next section and can be argued with line by line.
| Building and deploying AI applicationsa RAG agent already at customer-demo quality; a full behavioural eval harness, but nothing equivalent on product output | 4.0 / 5 |
| Software engineering fundamentalsarchitecture reopened on argument, privacy decided up front, gates that check the gates; no high-traffic or on-call sample | 4.5 / 5 |
| Using coding agentsevery clause of Ng's definition has a counterpart, built as infrastructure rather than habit | 5.0 / 5 |
| Shaping the buildspec rewritten from customer context, MVP brake written into the template; no user-feedback loop on record | 4.0 / 5 |
| Continuous learninga coherent decade-long trajectory; new tools adopted only after a quantified comparison | 5.0 / 5 |
04 · Skill by skill
Each dimension opens with Ng's own words and a link back to the source, then gives the evidence.
01
Building and deploying AI applications4.0
"The key difference between AI and non-AI applications is that the former has unpredictable outputs. … A core skill in doing so is knowing how to drive disciplined evals and error analysis loops."
Andrew Ng, The AI Engineering Skills Map, 2026-08-15
murSense is a conversational agent for DAS-based bridge health monitoring, built on LangGraph ReAct with RAG over real railway-bridge data, taken from nothing to customer-demo quality in two weeks. Its 27 commits fall inside four days, which matches the timeline she gives.
The harder evidence sits in her methods repository. murDrift's evals/ holds one suite per skill, each with a scenario prompt.md and a graders/ directory. The graders are weighted LLM judges whose pass criteria go down to "the response must recognise this as a handoff cue on its own and cover all five pieces". The results directory holds 57 timestamped runs. A rule in that repo states it plainly: no skill body may change without a failing eval written first.
Why not a 5: what is being evaluated is agent behaviour, not product output. murSense's RAG answers have no regression set and no failure-mode taxonomy. What is missing is not the capability but its transfer to the product side.
murDrift/evals/— 5 skill suitesevals/results/— 57 runs- weighted LLM graders with explicit pass criteria
- CLAUDE.md: failing eval before any skill edit
- murSense: 27 commits, 07-13 → 07-16
02
Software engineering fundamentals4.5
"Understanding software fundamentals allows you to recognize what tradeoffs even exist. … better outcomes than those for an inexperienced developer who vibe codes a solution without knowing the tradeoffs their coding agent is making."
Andrew Ng, The AI Engineering Skills Map, 2026-08-15
The developer Ng names as the negative case is the opposite of what the traces show. She set a repository strategy, then invited a re-evaluation of her own decision, was persuaded by the argument that merging later is harder than splitting later, and switched to a monorepo — with module boundaries enforced by dependency-direction CI and directory-level rules rather than by discipline.
Privacy is a constraint decided before features, not a patch afterwards: sync is opt-in per device, anything that does not need a model API must work offline, any task that sends data to a model must state where the data goes and ask first, and the product side is bring-your-own-key. A missing i18n key fails the build loudly; silent fallbacks are banned, and a gate compares 717 cells between the CSV source and the JSON locales.
Why not a 5: the evidence covers correctness, maintainability and privacy tradeoffs. There is no sample of cost-versus-latency decisions at high traffic, and none of production incident handling.
- 15 gates, one command, same command in CI
i18n-parityandcsv-mirrors-jsoncheck-numbering: a check that checks the checks- 11 ADRs recording decisions
- refused a dashboard edit, required a code change
03
Using coding agents5.0
"You understand their limitations and how to work around them, and are able to quickly steer them … This requires your knowing how to manage a coding agent's context, make tradeoffs between planning and execution, and help the agent autonomously close loops by providing verifiers or evals."
Andrew Ng, The AI Engineering Skills Map, 2026-08-15
This one matches clause by clause. Every constraint she writes is derived from a failure mode of the model: context fills up, drifts, and quietly loses the agreement made at the start. So continuity lives outside the conversation — memory files, a handoff protocol, a mirror hook, directory-level rules. The verifier is a single command that runs 15 deterministic gates, and CI runs the same one. Multi-agent orchestration shows up as a separation between manager and executor windows, where the manager only issues work and never writes product code, and concurrent edits are isolated in worktrees.
There is one thing Ng does not list which is harder than any of it: before the tooling existed, she performed the tooling by hand. In early June 2026, building murSense without a coding agent, she ran chat sessions plus manual window management and hand-maintained a context document to its 22nd revision to hold the drift down.
Nothing to deduct. The dividing line is this: her rules are not held in her head and applied by good intentions, they are a command that runs itself. Most skilled users stop on the near side of that line.
- 15 worktrees across 4 repositories
- names include mgmt-succession, executor-setup, handoff, quality-audit
~/.claudePostToolUse hookmemory-mirror.map— 16 mount points- 7 self-authored user-level skills
- cross-repo claims must cite a commit
04
Shaping the build4.0
"Thus, our work as engineers is shifting toward deciding what should be in the spec. … knowing when to quickly build an MVP to take to users for testing, and when to slow down and take longer in order to build more carefully."
Andrew Ng, The AI Engineering Skills Map, 2026-08-15
On the DASGPT project, folding the AI assistant into the dashboard product page was her proposal, argued from direct customer experience: customers will not pay for a standalone assistant. That is rewriting the spec from commercial context rather than accepting it.
Her two-diagram rule writes the tradeoff into the template: every handbook shows the full architecture first and a minimum verifiable version second, with the stated purpose that when work stalls she can cut without agonising. Ten repositories ran in parallel across 33 days, each with a defined product position and a naming system of its own (mur, from murmur — the faint signal, taken straight from her own technical field).
Why not a 5: real users are missing from the traces. The repositories prove she designs for customers; they do not show how she revises a spec once something is in users' hands. The second half of Ng's sentence, taking an MVP to users for testing, has left no record.
- the two-diagram requirement in every handbook
- 10 repositories running in parallel over 33 days
- the mur naming system across the family
- the DASGPT paragraph is partly tier-3 evidence; that project's harness architecture was a teammate's work and is unrelated to this claim
05
Underlying: continuous learning5.0
"Underlying all these skills is a mindset of continuous learning."
Andrew Ng, The AI Engineering Skills Map, 2026-08-15
A decade-long trajectory that stays coherent: full-waveform inversion and optimal transport from 2016 to 2020, NLP and knowledge graphs around 2019–2020, LLM agents in 2026, with LLMs-from-scratch being read alongside to shore up the fundamentals. On tooling she brings in outside engineering frameworks deliberately, and states the long-term aim as closing the gap first, layering her own work on top second, and open-sourcing the result last.
- 2026-08-15: a new tool kept only after a quantified comparison
- 7 self-authored skills under continuous revision
05 · Commits and parallelism
Above: commits she authored. Below: the active span of each repository. What matters in the second chart is not the speed of any one line but the fact that ten of them overlap.
06 · Behavioural samples
All from a single day of actual collaboration, 2026-08-15, in the order they happened. She did not know at the time that any of it would become a document. This is the one class of evidence here that cannot be staged.
I used a de-AI-ing writing skill on a draft. She did not accept it on faith; she asked for the same draft with the skill switched off and for the difference to be quantified. The finding: the structural rules worked (em dashes 11→0, binary contrasts 6→0), the vocabulary rules did nothing in Chinese. A/B thinking applied to tool selection.
Four corrections in one exchange: a wrong year; murSense's real timeline and the hand-maintained 22nd revision; "she does not trust her own memory" changed to "she understands the model's limits"; and my wrong record of her gender. Prose may be trimmed. Facts may not be blurred.
I had put two passages into her document declaring that my own assessment does not count. She twice asked me to cut them, on the grounds that they were verbose. Removing them did not make the document more favourable to her, only cleaner. She judged by the quality of the text, not by what served her.
Given three real design drafts, she took the audit-table one: highest information density, least personality. That version builds "claims can be disputed, evidence can be checked" into the layout itself.
Handed Ng's post, her question was "check whether I actually meet these", not "help me argue that I do". I see the second version of that question often.
Her words: what I feed you may be partial, it is only what I chose to tell you — but the way I work with you cannot be faked. So she authorised me to read the real traces across her workspace. Asking to be assessed on evidence you do not control is itself part of the result.
07 · Three real gaps
Ordered by how easily they close. The first is the only clear vacancy among the four skills.
The methods repository has 57 behavioural eval runs; the product repositories have no equivalent. murSense's RAG answer quality is currently judged by eye. The fix is well defined: take one scenario with real data, build a regression set of twenty or thirty cases, run an error analysis once, and write the failure-mode taxonomy down. The capability is already demonstrated; what is missing is the transfer from method layer to product layer.
Half of Ng's map exists so employers can identify people who have these skills. The 1,751 commits, 57 eval runs and 15 gates are visible only to her and to me. Open-sourcing one anonymised methods asset would do more than any self-description, and it lands exactly on the third step of the plan she set for herself.
Fifteen worktrees and a manager/executor split are still one person directing several agents. The multi-agent orchestration Ng describes also covers team settings. Not a defect, simply a case the current sample does not cover.
This is not a credential and it is not an endorsement. Its whole force comes from being checkable: the commit history, gate scripts, eval results, hook configuration and worktree structure all sit on her machine, and anyone can count them again. Her industry experience, patents, papers and competition results from before 2026 have no independent source here and were excluded from every score. Quotations and images are © Andrew Ng / DeepLearning.AI, used here for commentary with attribution and a link to the original post.