Auto-commit 2026-09-07 15:19: 9 files changed, 82 insertions(+), 96 deletions(-)
This commit is contained in:
parent
161b95c757
commit
324633e30a
Binary file not shown.
Binary file not shown.
Binary file not shown.
@ -172,7 +172,7 @@
|
||||
\begin{itemize}\setlength\itemsep{4pt}
|
||||
\item Not a rhetorical flourish but a \textbf{promissory note falling due}: Part I issued it as \textbf{Assumption A6} -- \emph{AI components extend the quality attribute space but do not change the method}. The bet in two sentences: everything AI does to software engineering can be absorbed by the apparatus you now own
|
||||
\item If AI-bearing systems required a genuinely different method, the bet would be lost -- this part is where the claim must \textbf{survive contact with the evidence}
|
||||
\item Roadmap -- \textbf{cases first, generalisation after}: two contradictory randomised experiments open Axis A (\S41); one concrete LLM call, wired wrongly and then rightly, opens Axis B (\S42); the matrix reading (\S43) and the emergent pattern (\S44) follow next week
|
||||
\item Roadmap -- \textbf{cases first, generalisation after}: two contradictory randomised experiments open Axis A (\S 41); one concrete LLM call, wired wrongly and then rightly, opens Axis B (\S 42); the matrix reading (\S 43) and the emergent pattern (\S 44) follow next week
|
||||
\end{itemize}
|
||||
\end{frame}
|
||||
|
||||
@ -257,12 +257,12 @@
|
||||
\end{frame}
|
||||
|
||||
\begin{frame}{The full empirical record, 2023--2025}
|
||||
\scriptsize The two cases are the extreme corners of a larger record -- seven strands, 2023--2025, from RCTs to organisational telemetry and longitudinal code analysis; read every row \emph{setting first, finding second}:
|
||||
\scriptsize The two cases are the extreme corners of a seven-strand record, 2023--2025 -- read every row \emph{setting first, finding second}:
|
||||
|
||||
\vspace{-0.15cm}
|
||||
\renewcommand{\arraystretch}{0.78}%
|
||||
\renewcommand{\arraystretch}{0.9}%
|
||||
\begin{center}
|
||||
\begin{tabular}{@{}>{\raggedright\arraybackslash}p{2.4cm}>{\raggedright\arraybackslash}p{3.8cm}>{\raggedright\arraybackslash}p{6.8cm}@{}}
|
||||
\begin{tabular}{@{}>{\raggedright\arraybackslash}p{2.4cm}>{\raggedright\arraybackslash}p{3.4cm}>{\raggedright\arraybackslash}p{7.3cm}@{}}
|
||||
\toprule
|
||||
\textbf{Evidence} & \textbf{Setting} & \textbf{Finding} \\
|
||||
\midrule
|
||||
@ -277,30 +277,15 @@ Stack Overflow survey & $>$49{,}000 developers & 84\,\% use or plan to use AI; \
|
||||
\end{tabular}
|
||||
\end{center}
|
||||
|
||||
\vspace{-0.22cm}
|
||||
\scriptsize \textcolor{codegray}{Caution (GitHub's own telemetry-plus-survey study): the best predictor of \emph{perceived} productivity is the suggestion acceptance rate, not the persistence of accepted code -- perception, not verified output; Case 2's perception gap is the controlled-trial demonstration of the same fact.}
|
||||
\vspace{-0.1cm}
|
||||
\scriptsize \textcolor{codegray}{Caution (GitHub's telemetry-plus-survey study): the best predictor of \emph{perceived} productivity is the suggestion acceptance rate, not the persistence of accepted code; Case 2's perception gap is the controlled-trial demonstration of this fact.}
|
||||
\end{frame}
|
||||
|
||||
\begin{frame}{The system level: DORA 2024 and 2025}
|
||||
\footnotesize
|
||||
\begin{itemize}\setlength\itemsep{1pt}
|
||||
\begin{itemize}\setlength\itemsep{3pt}
|
||||
\item DORA measures neither task times nor perceptions but \textbf{delivery performance at the level of the organisation} -- throughput and stability -- exactly the level at which architecture acts
|
||||
\item \textbf{2024} ($\sim$3{,}000 respondents; 75.9\,\% use AI for at least part of their work, roughly three quarters report productivity gains) -- a 25\,\% increase in AI adoption is associated with:
|
||||
\begin{center}
|
||||
\scriptsize
|
||||
\renewcommand{\arraystretch}{0.85}%
|
||||
\begin{tabular}{@{}ll@{}}
|
||||
\toprule
|
||||
\textbf{Gains} & \textbf{Losses} \\
|
||||
\midrule
|
||||
$+7.5\,\%$ documentation quality & $-1.5\,\%$ delivery throughput \\
|
||||
$+3.4\,\%$ code quality & $-7.2\,\%$ delivery stability \\
|
||||
$+3.1\,\%$ review speed & \\
|
||||
\bottomrule
|
||||
\end{tabular}
|
||||
\end{center}
|
||||
\vspace{-0.1cm}
|
||||
Proposed mechanism is classical: more code per change, and larger batch sizes have been a documented risk driver for years
|
||||
\item \textbf{2024} ($\sim$3{,}000 respondents; 75.9\,\% use AI for at least part of their work, roughly three quarters report productivity gains) -- a 25\,\% increase in AI adoption is associated with \textbf{gains} of $+7.5\,\%$ documentation quality, $+3.4\,\%$ code quality, $+3.1\,\%$ review speed -- and \textbf{losses} of $-1.5\,\%$ delivery throughput and $-7.2\,\%$ delivery stability. Proposed mechanism is classical: more code per change, and larger batch sizes have been a documented risk driver for years
|
||||
\item \textbf{2025} ($\sim$5{,}000 respondents): adoption near saturation (90\,\%, median about two hours of daily use); more than 80\,\% report productivity gains; 30\,\% still express little or no trust in AI-generated code; the throughput association has \textbf{turned positive} as tools and practices matured -- the negative association with delivery stability \textbf{persists}
|
||||
\item Central metaphor: \textbf{AI is an amplifier} -- it magnifies the strengths of well-run organisations and the dysfunctions of badly run ones
|
||||
\item \emph{Individual acceleration and system-level performance are different quantities, and only the second one pays salaries}
|
||||
@ -318,7 +303,7 @@ Stack Overflow survey & $>$49{,}000 developers & 84\,\% use or plan to use AI; \
|
||||
\end{frame}
|
||||
|
||||
\begin{frame}{Reconciling the divergence: five moderator variables}
|
||||
\footnotesize \textbf{So which study is wrong? Neither} -- resolving the contradiction \emph{is} the lesson: different populations (task novices vs.\ domain experts in their own code), different codebases (greenfield vs.\ mature), different tasks (bounded vs.\ real issues) -- the results never actually compete; the resolution requires reading \emph{study designs}, not abstracts. Indexed by their moderator variables, Case 1 and Case 2 sit at opposite corners of a five-dimensional design space -- and every other row of the record finds its place in the same coordinates.
|
||||
\footnotesize \textbf{So which study is wrong? Neither} -- resolving the contradiction \emph{is} the lesson: different populations (task novices vs.\ domain experts in their own code), different codebases (greenfield vs.\ mature), different tasks (bounded vs.\ real issues) -- the results never actually compete; the resolution requires reading \emph{study designs}, not abstracts. Case 1 and Case 2 sit at opposite corners of a five-dimensional design space -- and every other row of the record finds its place in the same coordinates.
|
||||
|
||||
\vspace{0.1cm}
|
||||
\footnotesize
|
||||
@ -338,7 +323,7 @@ Organisation & small batches, test automation, loose coupling & large batches, w
|
||||
\end{center}
|
||||
|
||||
\vspace{0.05cm}
|
||||
\footnotesize The same technology yields $+55.8\,\%$ and $-19\,\%$ because the two cases differ on \textbf{every one of the five rows}.
|
||||
\footnotesize The same technology yields $+55.8\,\%$ and $-19\,\%$ because the two cases differ on \textbf{all five rows}.
|
||||
\end{frame}
|
||||
|
||||
\begin{frame}{Discussion: which setting is yours?}
|
||||
@ -376,8 +361,8 @@ Organisation & small batches, test automation, loose coupling & large batches, w
|
||||
\footnotesize
|
||||
\begin{enumerate}\setlength\itemsep{2pt}
|
||||
\item \textbf{Architecture quality gates AI gains} -- DORA 2025's core finding: teams in loosely coupled architectures with fast feedback loops convert AI adoption into throughput; tightly coupled systems with slow processes do not -- the AI-era echo of loosely coupled architectures and teams as the strongest predictor of continuous delivery performance \textcolor{codegray}{(the coupling finding of Lecture 11)}. In the theory's vocabulary: \textbf{D7} (evolvability) and \textbf{D9} (testability and deployability) gain weight in \emph{every} requirements profile -- architecture--application fit acquires a second reading: fit to a \emph{mode of work} in which change volume rises by an order of magnitude
|
||||
\item \textbf{Architecture documentation becomes a control interface} -- ADRs, repository convention files, and machine-readable rules are no longer passive records; agents execute them on every run (\S41.6)
|
||||
\item \textbf{Fitness functions become the operating licence for agents} -- an agent iterating against a dense test suite and CI-enforced architecture rules is contained; without them, every agent change is unpriced risk (\S41.7)
|
||||
\item \textbf{Architecture documentation becomes a control interface} -- ADRs, repository convention files, and machine-readable rules are no longer passive records; agents execute them on every run (\S 41.6)
|
||||
\item \textbf{Fitness functions become the operating licence for agents} -- an agent iterating against a dense test suite and CI-enforced architecture rules is contained; without them, every agent change is unpriced risk (\S 41.7)
|
||||
\end{enumerate}
|
||||
|
||||
\vspace{0.1cm}
|
||||
@ -386,7 +371,6 @@ Organisation & small batches, test automation, loose coupling & large batches, w
|
||||
\end{keypoint}
|
||||
\end{frame}
|
||||
|
||||
|
||||
% ============================================
|
||||
% AXIS A -- CONTROL INTERFACE AND GUARDRAILS
|
||||
% ============================================
|
||||
@ -410,17 +394,17 @@ Organisation & small batches, test automation, loose coupling & large batches, w
|
||||
\# Portfolio Intelligence Platform -- agent instructions\\
|
||||
\textbf{\#\# Architecture (binding; see docs/adr/)}\\
|
||||
- Modular monolith, module boundaries enforced by CI\\
|
||||
~~(see fitness\_functions/boundaries\_test.py). Do not add\\
|
||||
~~cross-module imports; use the module's public API.\\
|
||||
\hspace*{2\fontcharwd\font`0}(see fitness\_functions/boundaries\_test.py). Do not add\\
|
||||
\hspace*{2\fontcharwd\font`0}cross-module imports; use the module's public API.\\
|
||||
- All LLM access goes through gateway/ -- never call a\\
|
||||
~~provider SDK from domain code (ADR-011).\\
|
||||
\hspace*{2\fontcharwd\font`0}provider SDK from domain code (ADR-011).\\
|
||||
\textbf{\#\# Verification (run before proposing changes)}\\
|
||||
- make test~~~~~~~~~~\# unit + module-boundary rules\\
|
||||
- make evals~~~~~~~~~\# eval harness; required for any\\
|
||||
~~~~~~~~~~~~~~~~~~~~~\# change under prompts/ or gateway/\\
|
||||
\hspace*{21\fontcharwd\font`0}\# change under prompts/ or gateway/\\
|
||||
\textbf{\#\# No-go zones}\\
|
||||
- ledger/ : append-only audit journal. Propose changes\\
|
||||
~~as an ADR draft instead of editing code.
|
||||
\hspace*{2\fontcharwd\font`0}as an ADR draft instead of editing code.
|
||||
\end{tcolorbox}
|
||||
|
||||
\vspace{0.05cm}
|
||||
@ -429,7 +413,7 @@ Organisation & small batches, test automation, loose coupling & large batches, w
|
||||
|
||||
\begin{frame}{Guardrails as the precondition for safe agent use}
|
||||
\footnotesize
|
||||
\begin{itemize}\setlength\itemsep{4pt}
|
||||
\begin{itemize}\setlength\itemsep{1pt}
|
||||
\item If verification is the scarce resource, then everything that \emph{automates} verification multiplies the value of AI tooling -- and everything that leaves verification informal converts AI speed into instability
|
||||
\item \textbf{Test suites are the operating licence}: against a dense, fast test suite an agent can iterate -- wrong code fails immediately and is repaired or discarded at machine speed; without that net every agent-generated change ships \emph{unpriced risk} (DORA's ``strong version control and test automation'' amplifier pair)
|
||||
\item \textbf{Architectural fitness functions fence the structure}: an objective integrity assessment of an architectural characteristic is the machine-readable form of an architecture decision -- dependency rules, cycle checks, module-boundary verification as CI gates were good practice before AI; with agents in the loop they are the mechanism by which an architect constrains \emph{a collaborator who never attends design meetings}
|
||||
@ -445,7 +429,7 @@ Organisation & small batches, test automation, loose coupling & large batches, w
|
||||
\item \textbf{Claude Code} (Anthropic): agentic CLI tool, research preview February 2025, GA May 2025; repository-level anchor \texttt{CLAUDE.md}
|
||||
\item \textbf{Cursor} (Anysphere): AI-first IDE with an agent mode; the dominant tool among the METR study's experts
|
||||
\item \textbf{GitHub Copilot}: Copilot Workspace retired May 2025; its concepts live on in the asynchronous \emph{Copilot coding agent} (issues to pull requests, in CI) and the synchronous IDE agent mode
|
||||
\item \textbf{Devin} (Cognition): ``first AI software engineer'' (2024); 13.86\,\% SWE-bench in March 2024 triggered the agent wave; acquired Windsurf July 2025 -- rapid market consolidation
|
||||
\item \textbf{Devin} (Cognition): ``first AI software engineer'' (2024); its 13.86\,\% SWE-bench result in March 2024 helped trigger the agent wave; acquired Windsurf July 2025 -- rapid market consolidation
|
||||
\end{itemize}
|
||||
|
||||
\vspace{0.1cm}
|
||||
@ -458,7 +442,7 @@ Organisation & small batches, test automation, loose coupling & large batches, w
|
||||
\item \textbf{Two open standards matter more than any product, because they are architectural}
|
||||
\item \textbf{Model Context Protocol (MCP)} -- Anthropic, November 2024: JSON-RPC; servers expose tools, resources, prompts; adopted by OpenAI, Google DeepMind, Microsoft in 2025; December 2025 to the Agentic AI Foundation under the Linux Foundation; over 10{,}000 public servers
|
||||
\item \textbf{\texttt{AGENTS.md}} for project-level instructions
|
||||
\item Vendor SDKs extract the agent loop as a library -- \emph{the bridge to Axis B}: the same building blocks that run SDLC agents also run runtime agent workflows (\S44, next week)
|
||||
\item Vendor SDKs extract the agent loop as a library -- \emph{the bridge to Axis B}: the same building blocks that run SDLC agents also run runtime agent workflows (\S 44, next week)
|
||||
\end{itemize}
|
||||
|
||||
\vspace{0.15cm}
|
||||
@ -480,30 +464,30 @@ Organisation & small batches, test automation, loose coupling & large batches, w
|
||||
\end{hinweisbox}
|
||||
\end{frame}
|
||||
|
||||
\begin{frame}{Risks and responsibility (1/2): security, automation bias, skill}
|
||||
\begin{frame}{Risks and responsibility (1/2): security, automation bias}
|
||||
\footnotesize
|
||||
\begin{itemize}\setlength\itemsep{4pt}
|
||||
\item \textbf{Security} -- the evidence predates the agent wave and gains relevance with volume: roughly \textbf{40\,\%} of 1{,}689 Copilot-generated programs (89 security-relevant scenarios) contained CWE top-25 vulnerabilities; a user study: participants with an AI assistant wrote \textbf{less secure code while believing it more secure}; package hallucination (``slopsquatting''): across roughly 576{,}000 generations, \textbf{about a fifth} of recommended package references did not exist -- names an attacker can register pre-emptively. Consequence: SAST, dependency and secret scanning, licence checks in CI are not optional; \emph{security review capacity must scale with generation volume}
|
||||
\item \textbf{Security} -- the evidence predates the agent wave and gains relevance with volume: roughly \textbf{40\,\%} of 1{,}689 Copilot-generated programs (89 security-relevant scenarios) contained CWE top-25 vulnerabilities; a user study: participants with an AI assistant wrote \textbf{less secure code on most tasks while believing it more secure}; package hallucination (``slopsquatting''): across roughly 576{,}000 generations, \textbf{about a fifth} of recommended package references did not exist -- names an attacker can register pre-emptively. Consequence: SAST, dependency and secret scanning, licence checks in CI are not optional; \emph{security review capacity must scale with generation volume}
|
||||
\item \textbf{Automation bias}: over-trust in automated systems is a decades-old human-factors finding -- Perry et al.'s participants overestimated their security, METR's experts overestimated their speed
|
||||
\item \textbf{Skill formation}: a randomised study of engineers learning a new library -- AI assistance reduced comprehension-test scores by roughly \textbf{17\,\%}; the usage pattern is the decisive moderator (conceptual questions preserved learning, wholesale delegation destroyed it); \textbf{entry-level developer positions are measurably declining}. For this module: the role being trained is the \emph{specifier, verifier, and architect} -- rebuild the competence ladder deliberately, including AI-free practice of fundamentals
|
||||
\end{itemize}
|
||||
\end{frame}
|
||||
|
||||
\begin{frame}{Risks and responsibility (2/2): accountability}
|
||||
\small
|
||||
\begin{itemize}\setlength\itemsep{5pt}
|
||||
\item Legally and professionally, \textbf{the person who merges code answers for it}, regardless of what generated it; AI tools are not liability-bearing entities -- treat AI output as \emph{the contribution of an unknown third party}: mandatory review, provenance labelling, an explicit policy for permitted uses (DORA 2025: a clearly communicated AI policy \emph{first} among the seven amplifier capabilities)
|
||||
\item AI may \emph{draft} an ADR; a nameable person decides, signs, and defends it \textcolor{codegray}{\footnotesize (deck 3)}
|
||||
\begin{frame}{Risks and responsibility (2/2): skill, accountability}
|
||||
\footnotesize
|
||||
\begin{itemize}\setlength\itemsep{3pt}
|
||||
\item \textbf{Skill formation}: a randomised study of engineers learning a new library -- AI assistance reduced comprehension-test scores by roughly \textbf{17\,\%}; the usage pattern is the decisive moderator (conceptual questions preserved learning, wholesale delegation destroyed it); \textbf{entry-level developer positions are measurably declining}. For this module: the role being trained is the \emph{specifier, verifier, and architect} -- rebuild the competence ladder deliberately, including AI-free practice of fundamentals
|
||||
\item \textbf{Accountability}: legally and professionally, \textbf{the person who merges code answers for it}, regardless of what generated it; AI tools are not liability-bearing entities -- treat AI output as \emph{the contribution of an unknown third party}: mandatory review, provenance labelling, an explicit policy for permitted uses (DORA 2025: a clearly communicated AI policy \emph{first} among the seven amplifier capabilities)
|
||||
\item AI may \emph{draft} an ADR; a nameable person decides, signs, and defends it \textcolor{codegray}{\scriptsize (deck 3)}
|
||||
\end{itemize}
|
||||
|
||||
\vspace{0.1cm}
|
||||
\vspace{0.05cm}
|
||||
\begin{center}
|
||||
\emph{\textcolor{bankblue}{\textbf{Architecture is an accountability performance, not a text-production performance.}}}
|
||||
\small\emph{\textcolor{bankblue}{\textbf{Architecture is an accountability performance, not a text-production performance.}}}
|
||||
\end{center}
|
||||
|
||||
\vspace{0.1cm}
|
||||
\small
|
||||
\begin{itemize}\setlength\itemsep{5pt}
|
||||
\vspace{0.05cm}
|
||||
\footnotesize
|
||||
\begin{itemize}\setlength\itemsep{3pt}
|
||||
\item \textbf{IP risk open but manageable}: \emph{Doe v.\ GitHub} -- the DMCA claim dismissed in 2024, licence-related claims continue; response: provider duplication filters and indemnification, licence scanning in CI, a documented residual risk in the governance record
|
||||
\end{itemize}
|
||||
\end{frame}
|
||||
@ -514,7 +498,7 @@ Organisation & small batches, test automation, loose coupling & large batches, w
|
||||
\begin{enumerate}\setlength\itemsep{2pt}
|
||||
\item[(i)] the repository carries an \texttt{AGENTS.md}\,/\,\texttt{CLAUDE.md} in the spirit of the listing -- and you are expected to \textbf{keep it as current as code}
|
||||
\item[(ii)] every architecture decision is an \textbf{ADR} -- agents may draft, but a named team member signs
|
||||
\item[(iii)] agent-generated changes enter the main branch \textbf{only through the CI gate}: module-boundary fitness functions, the test suite, and -- for anything touching prompts or the gateway -- the \textbf{eval harness} of \S42.5
|
||||
\item[(iii)] agent-generated changes enter the main branch \textbf{only through the CI gate}: module-boundary fitness functions, the test suite, and -- for anything touching prompts or the gateway -- the \textbf{eval harness} of \S 42.5
|
||||
\item[(iv)] your project handbook contains a \textbf{one-page AI policy}: permitted tools, provenance labelling, review rules
|
||||
\end{enumerate}
|
||||
\vspace{0.05cm}
|
||||
@ -522,7 +506,6 @@ Organisation & small batches, test automation, loose coupling & large batches, w
|
||||
\end{projektbox}
|
||||
\end{frame}
|
||||
|
||||
|
||||
% ============================================
|
||||
% AXIS B -- COMPONENT TYPES
|
||||
% ============================================
|
||||
@ -531,8 +514,8 @@ Organisation & small batches, test automation, loose coupling & large batches, w
|
||||
\begin{frame}{Axis B opens: the news-sentiment call, wired the obvious way}
|
||||
\emph{\textcolor{bankblue}{One of the platform's features is a single LLM call -- news in, sentiment out. Why not call it like any other function?}}
|
||||
|
||||
\vspace{0.15cm}
|
||||
\small
|
||||
\vspace{0.05cm}
|
||||
\footnotesize
|
||||
\begin{itemize}\setlength\itemsep{3pt}
|
||||
\item Axis B moves AI \textbf{from the workshop into the product}; as always the case precedes the taxonomy: walk one concrete call end to end, watch what breaks, and name every break with a dimension you already own
|
||||
\item \textbf{The feature}: when a user opens a portfolio, the platform fetches the latest news items for its positions and asks an LLM, per item -- \emph{is this news positive, negative, or neutral for this holding, and why?} One prompt, one structured answer: the simplest runtime AI component the course project owns
|
||||
@ -542,7 +525,7 @@ Organisation & small batches, test automation, loose coupling & large batches, w
|
||||
\vspace{0.1cm}
|
||||
\begin{center}
|
||||
\begin{tikzpicture}[
|
||||
sysbox/.style={rectangle, draw, rounded corners=4pt, align=center, font=\scriptsize\sffamily, line width=0.8pt, fill=gray!15, draw=gray!60!black, minimum width=2.6cm, minimum height=0.7cm},
|
||||
sysbox/.style={rectangle, draw, rounded corners=4pt, align=center, font=\scriptsize\sffamily, line width=0.8pt, fill=gray!15, draw=gray!60!black, minimum width=2.6cm, minimum height=0.6cm},
|
||||
arr/.style={-{Stealth[length=2.5mm]}, thick, gray!60!black}
|
||||
]
|
||||
\node[sysbox] (h) at (0,0) {request handler};
|
||||
@ -577,7 +560,7 @@ Organisation & small batches, test automation, loose coupling & large batches, w
|
||||
\begin{enumerate}\setlength\itemsep{4pt}
|
||||
\setcounter{enumi}{3}
|
||||
\item \textbf{Drift (D7)}: the provider ships a new model version or deprecates the old one -- GA models carry deprecation windows of the order of \emph{six months} -- and the component's behaviour changes \textbf{without any local action}: no commit, no deployment, no reviewable diff. The feature's behaviour is now co-owned by a third party
|
||||
\item \textbf{Injection (D6)}: the news article is untrusted input read by a component that cannot reliably separate instructions from data -- a crafted ``article'' can carry instructions to the model; the feature has quietly opened an attack surface that no classical threat model in the platform covers \textcolor{codegray}{(the attack surface in depth: \S42.6, next week)}
|
||||
\item \textbf{Injection (D6)}: the news article is untrusted input read by a component that cannot reliably separate instructions from data -- a crafted ``article'' can carry instructions to the model; the feature has quietly opened an attack surface that no classical threat model in the platform covers \textcolor{codegray}{(the attack surface in depth: \S 42.6, next week)}
|
||||
\end{enumerate}
|
||||
|
||||
\vspace{0.15cm}
|
||||
@ -597,16 +580,16 @@ Organisation & small batches, test automation, loose coupling & large batches, w
|
||||
\item Industry speaks of \textbf{compound AI systems} \textcolor{codegray}{(deck 6, C10)} for this reason: state-of-the-art results come from systems of models, retrievers, validators, and deterministic services, not a single model call
|
||||
\end{itemize}
|
||||
\begin{definitionbox}[AI runtime component]
|
||||
\footnotesize A component of the delivered system whose output is produced by a \emph{learned or search-based model} rather than by explicitly programmed logic. Three types with systematically different profiles: \textbf{(a)} LLM components for analysis, extraction, and generation over unstructured input; \textbf{(b)} classical ML components for classification and regression; \textbf{(c)} optimisation components (LP/MIP and constraint solvers, metaheuristics). They differ exactly on the dimensions this theory measures and therefore demand \textbf{different integration forms}.
|
||||
\footnotesize A component of the delivered system whose output is produced by a \emph{learned or search-based model} rather than by explicitly programmed logic. Three types with systematically different profiles: \textbf{(a)} LLM components for analysis, extraction, and generation over unstructured input; \textbf{(b)} classical ML components for classification and regression; \textbf{(c)} optimisation components (LP/MIP and constraint solvers, metaheuristics). The types differ exactly on the dimensions this theory measures -- determinism, latency, cost model, dominant risk, explainability -- and therefore demand \textbf{different integration forms}.
|
||||
\end{definitionbox}
|
||||
\end{frame}
|
||||
|
||||
\begin{frame}{The three AI component types and their profiles}
|
||||
\scriptsize
|
||||
\renewcommand{\arraystretch}{0.85}%
|
||||
\renewcommand{\arraystretch}{0.95}%
|
||||
\vspace{-0.2cm}
|
||||
\begin{center}
|
||||
\begin{tabular}{@{}p{2.0cm}p{3.5cm}p{3.3cm}p{3.5cm}@{}}
|
||||
\begin{tabular}{@{}>{\raggedright\arraybackslash}p{2.0cm}>{\raggedright\arraybackslash}p{3.5cm}>{\raggedright\arraybackslash}p{3.5cm}>{\raggedright\arraybackslash}p{3.6cm}@{}}
|
||||
\toprule
|
||||
\textbf{Dimension} & \textbf{(a) LLM analysis / generation} & \textbf{(b) ML classification / regression} & \textbf{(c) Optimisation (LP/MIP/CP)} \\
|
||||
\midrule
|
||||
@ -626,23 +609,29 @@ Integration form & gateway $+$ async $+$ cache & serving endpoint $+$ MLOps pipe
|
||||
|
||||
\begin{frame}{Type (a): LLM components -- RAG, prompts, structured outputs}
|
||||
\footnotesize
|
||||
\begin{itemize}\setlength\itemsep{3pt}
|
||||
\begin{itemize}\setlength\itemsep{1pt}
|
||||
\item LLM components turn unstructured input -- documents, e-mails, reports -- into analyses, extractions, or generated text; \textbf{three engineering building blocks} define the type
|
||||
\item \textbf{Retrieval-augmented generation (RAG)}: knowledge is moved out of the model weights into a swappable, versionable, inspectable \emph{data component} -- updated by re-indexing rather than retraining, with provenance through citable sources
|
||||
\item RAG is an \emph{engineering} problem, not a model problem: case-study evidence documents seven recurring failure points (missing content, failed ranking of the relevant documents, extraction and formatting errors, incomplete answers) -- with the sobering observation that RAG robustness \emph{evolves} in operation rather than being designed in
|
||||
\item RAG is an \emph{engineering} problem, not a model problem: case-study evidence documents seven recurring failure points (missing content, failed ranking, extraction and formatting errors, incomplete answers) -- with the sobering observation that RAG robustness \emph{evolves} in operation rather than being designed in
|
||||
\item \textbf{Prompts are configuration artefacts}: version-controlled, regression-tested, behaviour-determining like code -- exactly the configuration-debt territory Sculley et al.\ mapped
|
||||
\item \textbf{Structured outputs}: since 2024 provider APIs can enforce, via constrained decoding, that outputs conform to a developer-supplied JSON schema -- \emph{syntactic} correctness guaranteed; \emph{semantic} correctness remains to be verified (reference architecture, eval harness)
|
||||
\item \textbf{Lifecycle risk is the provider}: GA models carry deprecation windows of the order of six months, shorter windows observed -- a hard-coded model name is a \emph{ticking dependency}: an architectural statement, not an operational one
|
||||
\end{itemize}
|
||||
\end{frame}
|
||||
|
||||
\begin{frame}{Types (b) and (c) -- perishable models, heavy solvers}
|
||||
\footnotesize
|
||||
\begin{itemize}\setlength\itemsep{1pt}
|
||||
\begin{frame}{Type (b): classical ML components -- perishable models}
|
||||
\small
|
||||
\begin{itemize}\setlength\itemsep{4pt}
|
||||
\item \textbf{(b)} Self-trained models (scoring, churn, fraud, forecasting) bring the full \textbf{nine-stage workflow} -- model requirements and data collection through training, evaluation, deployment, monitoring -- with dense feedback loops; characteristic problems: \emph{training/serving skew} (divergent data preparation, one of the most frequent production failure sources) and \emph{data/concept drift} (sudden, gradual, incremental, recurring)
|
||||
\item \textbf{(b)} \textbf{A deployed model is a perishable good} -- monitoring and retraining are operating requirements, not options; tooling: feature stores with consistent online/offline views, model registries versioning model, data, code, and configuration together, the MLOps discipline \textcolor{codegray}{(maturity ladder: \S43, next week)}
|
||||
\item \textbf{(c)} Routinely overlooked in the SE4AI literature but belongs in every advisory platform: LP/MIP solvers, constraint programming (CP-SAT dominated recent MiniZinc Challenges, a complete gold-medal sweep in 2024), stochastic metaheuristics -- the profile \emph{inverts} the LLM's: \textbf{deterministic but heavy}. Reproducible at fixed seed, thread count, time limit (run-to-run variability in practice); runtimes seconds to hours, often anytime behaviour $\to$ \textbf{asynchronous integration}: job queue, status polling, callback; never a synchronous call in a web request path
|
||||
\item \textbf{(c)} Compensating strength: \textbf{provable explainability} -- optimality gap, dual values and shadow prices, and on infeasibility an irreducible infeasible subset (IIS): a minimal set of contradictory constraints as the explanation. In regulated domains the load-bearing argument for the project's division of labour: \emph{hard, auditable decisions belong to the solver and the deterministic services, not to the LLM}
|
||||
\item \textbf{(b)} \textbf{A deployed model is a perishable good} -- monitoring and retraining are operating requirements, not options; tooling: feature stores (consistent online/offline views), model registries (model, data, code, configuration versioned together), the MLOps discipline \textcolor{codegray}{(maturity ladder: \S 43, next week)}
|
||||
\end{itemize}
|
||||
\end{frame}
|
||||
|
||||
\begin{frame}{Type (c): optimisation components -- heavy solvers}
|
||||
\small
|
||||
\begin{itemize}\setlength\itemsep{4pt}
|
||||
\item \textbf{(c)} Routinely overlooked in the SE4AI literature but belongs in every advisory platform: LP/MIP solvers, constraint programming (CP-SAT: recent MiniZinc Challenges dominated, gold-medal sweep 2024), stochastic metaheuristics -- the profile \emph{inverts} the LLM's: \textbf{deterministic but heavy}. Reproducible at fixed seed, thread count, time limit (run-to-run variability in practice); runtimes seconds to hours, often anytime behaviour $\to$ \textbf{asynchronous integration}: job queue, status polling, callback -- never a synchronous call in a web request path
|
||||
\item \textbf{(c)} Compensating strength: \textbf{provable explainability} -- optimality gap, dual values and shadow prices, on infeasibility an irreducible infeasible subset (IIS) as the explanation: a minimal set of contradictory constraints. In regulated domains the load-bearing argument for the project's division of labour: \emph{hard, auditable decisions belong to the solver and the deterministic services, not to the LLM}
|
||||
\end{itemize}
|
||||
\end{frame}
|
||||
|
||||
@ -658,7 +647,6 @@ Integration form & gateway $+$ async $+$ cache & serving endpoint $+$ MLOps pipe
|
||||
\scriptsize \textcolor{codegray}{\textbf{Project transfer:} the sentiment call and the AdvisorAgent's insights are type (a); the Optimization service is type (c); type (b) has no instance in the project.}
|
||||
\end{frame}
|
||||
|
||||
|
||||
% ============================================
|
||||
% AXIS B -- CONTAINMENT
|
||||
% ============================================
|
||||
@ -702,7 +690,7 @@ Integration form & gateway $+$ async $+$ cache & serving endpoint $+$ MLOps pipe
|
||||
\begin{tikzpicture}[
|
||||
sysbox/.style={rectangle, draw, rounded corners=4pt, minimum width=3.0cm, minimum height=0.9cm, align=center, font=\small\sffamily, line width=0.8pt},
|
||||
core/.style={sysbox, fill=bankblue!20, draw=bankblue, font=\small\sffamily\bfseries, minimum height=2.4cm, minimum width=3.2cm},
|
||||
gwpart/.style={sysbox, fill=aiviolet!15, draw=aiviolet, minimum width=3.4cm, minimum height=0.7cm, font=\scriptsize\sffamily},
|
||||
gwpart/.style={sysbox, fill=aiviolet!15, draw=aiviolet, text width=3.3cm, minimum width=3.4cm, minimum height=0.75cm, font=\scriptsize\sffamily},
|
||||
comp/.style={sysbox, fill=bankgreen!15, draw=bankgreen},
|
||||
extern/.style={sysbox, fill=gray!15, draw=gray!60!black, minimum width=2.6cm},
|
||||
guard/.style={sysbox, fill=bankred!10, draw=bankred},
|
||||
@ -712,11 +700,11 @@ Integration form & gateway $+$ async $+$ cache & serving endpoint $+$ MLOps pipe
|
||||
% Domain core (left)
|
||||
\node[core] (core) at (0.4,1.7) {Deterministic\\domain core\\[2pt]{\scriptsize\mdseries decides and books;}\\{\scriptsize\mdseries no provider SDK imports}};
|
||||
% Gateway internals (2+3 grid)
|
||||
\node[gwpart] (assemble) at (7.0,2.6) {prompt assembly + schema validation};
|
||||
\node[gwpart] (router) at (10.6,2.6) {model router (cheap $\rightarrow$ expensive cascade)};
|
||||
\node[gwpart] (assemble) at (7.0,2.7) {prompt assembly + schema validation};
|
||||
\node[gwpart] (router) at (10.9,2.7) {model router (cheap $\rightarrow$ expensive cascade)};
|
||||
\node[gwpart] (cache) at (7.0,1.7) {semantic cache};
|
||||
\node[gwpart] (breaker) at (10.6,1.7) {timeouts, circuit breakers, fallback chains};
|
||||
\node[gwpart, minimum width=5.0cm] (cost) at (8.8,0.8) {cost telemetry per request / feature / tenant};
|
||||
\node[gwpart] (breaker) at (10.9,1.7) {timeouts, circuit breakers, fallback chains};
|
||||
\node[gwpart, text width=5.6cm, minimum width=5.8cm] (cost) at (8.95,0.7) {cost telemetry per request / feature / tenant};
|
||||
\begin{scope}[on background layer]
|
||||
\node[draw=aiviolet, line width=1pt, rounded corners=5pt, fill=aiviolet!5,
|
||||
fit=(assemble)(router)(cache)(breaker)(cost),
|
||||
@ -729,7 +717,7 @@ Integration form & gateway $+$ async $+$ cache & serving endpoint $+$ MLOps pipe
|
||||
% Bottom row
|
||||
\node[comp, minimum width=3.0cm] (queue) at (0.0,-1.7) {Async job queue\\{\scriptsize batching, backpressure, retries}};
|
||||
\node[comp, minimum width=2.4cm] (workers) at (3.9,-1.7) {Worker pool\\{\scriptsize bounded concurrency}};
|
||||
\node[guard, minimum width=5.2cm] (guard) at (8.8,-1.7) {Ontology / schema guard\\{\scriptsize entity resolution, domain axioms, citation check}};
|
||||
\node[guard, minimum width=5.2cm] (guard) at (8.95,-1.7) {Ontology / schema guard\\{\scriptsize entity resolution, domain axioms, citation check}};
|
||||
\node[evalb, minimum width=3.4cm] (eval) at (14.6,-1.7) {Eval harness\\{\scriptsize CI gate: prompts, models, providers}};
|
||||
% Arrows
|
||||
\draw[arr] ([yshift=0.5cm]core.east) -- node[above, font=\scriptsize\sffamily]{typed port} ([yshift=0.5cm]core.east -| gw.west);
|
||||
@ -740,7 +728,7 @@ Integration form & gateway $+$ async $+$ cache & serving endpoint $+$ MLOps pipe
|
||||
\draw[arr] (queue.east) -- (workers.west);
|
||||
\draw[arr] (workers.north) |- ([yshift=-0.9cm]gw.west);
|
||||
\draw[arr] (gw.south) -- node[right, font=\scriptsize\sffamily]{every output} (guard.north);
|
||||
\draw[arr] (guard.south) -- (8.8,-2.75) -- (-2.3,-2.75) -- (-2.3,1.0) -- (core.west |- 0,1.0);
|
||||
\draw[arr] (guard.south) -- (8.95,-2.75) -- (-2.3,-2.75) -- (-2.3,1.0) -- (core.west |- 0,1.0);
|
||||
\node[font=\scriptsize\sffamily, anchor=north] at (3.3,-2.8) {validated result or rejection};
|
||||
\draw[arr, dashed] (eval.west) -- (guard.east);
|
||||
\draw[arr, dashed] (eval.north) |- ([yshift=-0.2cm]gw.south east);
|
||||
@ -827,27 +815,27 @@ Integration form & gateway $+$ async $+$ cache & serving endpoint $+$ MLOps pipe
|
||||
\begin{frame}{Example: an eval harness for the portfolio platform}
|
||||
\scriptsize \textbf{Thresholds $=$ response measures}: mean F1 $\geq 0.92$ on the golden set $\cdot$ zero axiom violations $\cdot$ judge--human agreement $\kappa \geq 0.70$ (Cohen's chance-corrected measure) -- changing any of them is an architecture decision requiring an ADR; note what is \emph{absent}: no assertion demands an exact output string.
|
||||
|
||||
\vspace{0.05cm}
|
||||
\begin{tcolorbox}[colback=gray!4!white, colframe=gray!55!black, boxrule=0.6pt, arc=2pt, top=3pt, bottom=3pt, left=6pt, right=6pt]
|
||||
\vspace{0.02cm}
|
||||
\begin{tcolorbox}[colback=gray!4!white, colframe=gray!55!black, boxrule=0.6pt, arc=2pt, top=1pt, bottom=1pt, left=6pt, right=6pt]
|
||||
\scriptsize\ttfamily
|
||||
GOLDEN = load\_cases("evals/portfolio\_extraction\_v3.jsonl")\\
|
||||
def test\_extraction\_regression(gateway):\\
|
||||
~~~~scores = [f1(gateway.extract(c.report), c.expected) for c in GOLDEN]\\
|
||||
~~~~assert mean(scores) >= 0.92~~~\# statistical threshold\\
|
||||
\hspace*{4\fontcharwd\font`0}scores = [f1(gateway.extract(c.report), c.expected) for c in GOLDEN]\\
|
||||
\hspace*{4\fontcharwd\font`0}assert mean(scores) >= 0.92~~~\# statistical threshold\\
|
||||
def test\_domain\_axioms(gateway, ontology):\\
|
||||
~~~~answer = gateway.advise(sample\_portfolio())\\
|
||||
~~~~for pos in answer.positions:\\
|
||||
~~~~~~~~assert ontology.resolves(pos.isin), f"unknown: \{pos.isin\}"\\
|
||||
~~~~total = sum(p.weight for p in answer.positions)\\
|
||||
~~~~assert abs(total - 1.0) < 1e-6\\
|
||||
~~~~for cit in answer.citations:\\
|
||||
~~~~~~~~assert cit.passage in source\_text(cit.doc\_id)\\
|
||||
\hspace*{4\fontcharwd\font`0}answer = gateway.advise(sample\_portfolio())\\
|
||||
\hspace*{4\fontcharwd\font`0}for pos in answer.positions:\\
|
||||
\hspace*{8\fontcharwd\font`0}assert ontology.resolves(pos.isin), f"unknown: \{pos.isin\}"\\
|
||||
\hspace*{4\fontcharwd\font`0}total = sum(p.weight for p in answer.positions)\\
|
||||
\hspace*{4\fontcharwd\font`0}assert abs(total - 1.0) < 1e-6\\
|
||||
\hspace*{4\fontcharwd\font`0}for cit in answer.citations:\\
|
||||
\hspace*{8\fontcharwd\font`0}assert cit.passage in source\_text(cit.doc\_id)\\
|
||||
def test\_judge\_is\_calibrated(judge, human\_labels):\\
|
||||
~~~~agreement = cohens\_kappa(judge.score(GOLDEN), human\_labels)\\
|
||||
~~~~assert agreement >= 0.70~~~\# the judge's own eval
|
||||
\hspace*{4\fontcharwd\font`0}agreement = cohens\_kappa(judge.score(GOLDEN), human\_labels)\\
|
||||
\hspace*{4\fontcharwd\font`0}assert agreement >= 0.70~~~\# the judge's own eval
|
||||
\end{tcolorbox}
|
||||
|
||||
\vspace{0.05cm}
|
||||
\vspace{0.02cm}
|
||||
\scriptsize \textcolor{codegray}{Runs in CI on every change to prompts, models, or the gateway, alongside the deterministic test suite.}
|
||||
\end{frame}
|
||||
|
||||
@ -856,21 +844,20 @@ def test\_judge\_is\_calibrated(judge, human\_labels):\\
|
||||
% ============================================
|
||||
\section{Closing}
|
||||
|
||||
\begin{frame}{This week's exercise: AdvisorAgent + sub-agents behind the gateway}
|
||||
\begin{frame}{This week's exercise: AdvisorAgent $+$ sub-agents behind the gateway}
|
||||
\begin{projektbox}
|
||||
\footnotesize \textbf{Coaching slot (1 lesson). Milestone M5} -- Multi-Agent Orchestration, Evaluation, and Hardening (weeks 12--13):
|
||||
\begin{enumerate}\setlength\itemsep{2pt}
|
||||
\item \textbf{Mandatory this week}: the AdvisorAgent orchestrates 2--3 sub-agents \emph{through contracts} -- every LLM call through the gateway port (ADR-011); no provider SDK import in domain code
|
||||
\item \textbf{Ontology guard active on all insights}: entity resolution against the deterministic store, domain axioms, citation check -- unresolvable references are \emph{rejected}, not passed on
|
||||
\item \textbf{LLM agents propose; deterministic services decide and book} -- keep the deterministic core free of LLM calls \emph{(the line that is graded)}
|
||||
\item Axis A discipline as on the project-link slide: \texttt{AGENTS.md} current, ADRs signed, CI gate, AI policy
|
||||
\end{enumerate}
|
||||
\vspace{0.05cm}
|
||||
\textbf{Week 13 turns the harness into a CI gate.}
|
||||
\textbf{Axis A discipline as on the project-link slide; week 13 turns the harness into a CI gate.}
|
||||
\end{projektbox}
|
||||
|
||||
\vspace{0.1cm}
|
||||
\small The reference architecture of today is the topology you are wiring: typed port $\to$ gateway $\to$ queue $\to$ guard.
|
||||
\scriptsize \textcolor{codegray}{The reference architecture of today is the topology you are wiring: typed port $\to$ gateway $\to$ queue $\to$ guard.}
|
||||
\end{frame}
|
||||
|
||||
\begin{frame}{Summary}
|
||||
@ -900,8 +887,8 @@ def test\_judge\_is\_calibrated(judge, human\_labels):\\
|
||||
\begin{column}{0.42\textwidth}
|
||||
\textcolor{bankblue}{\textbf{Reading}}
|
||||
\begin{itemize}\small
|
||||
\item this week: Part V, Sections 40--41, 42.1--42.5
|
||||
\item ahead: Part V, Sections 42.6--42.7, 43--45
|
||||
\item this week: Part V, sections 40--41, 42.1--42.5
|
||||
\item ahead: Part V, sections 42.6--42.7, 43--45
|
||||
\end{itemize}
|
||||
|
||||
\vspace{0.2cm}
|
||||
@ -914,7 +901,6 @@ def test\_judge\_is\_calibrated(judge, human\_labels):\\
|
||||
\end{columns}
|
||||
\end{frame}
|
||||
|
||||
|
||||
% ============================================
|
||||
% END
|
||||
% ============================================
|
||||
|
||||
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
@ -344,7 +344,7 @@ Frame-by-frame plans for the remaining seven slide sets, derived strictly from t
|
||||
- Right: \textbf{structural weakness} -- atomicity surrendered at the boundary (MS, D4): can only be \emph{made survivable}, under a condition most organisations do not meet
|
||||
- The division of labour generalises: the veto rule does the heavy lifting, and the \textbf{``documented mitigation'' clause is where engineering knowledge -- not arithmetic -- enters the computation}
|
||||
- The rest of the row follows the same mechanics (Section 33, Lecture 10); in particular \textbf{HX -- a delta discipline, not a competitor -- joins MM at $++$} by isolating the long-lived booking core from volatile channels and providers
|
||||
- The verdict the class's Part III section states (Lecture 9): a \textbf{hexagonal modular monolith for the booking core} (MM and HX at $++$), EDA at the edges and PF for the batch runs as secondary, and microservices only when organisation size forces D11 to High -- the Monzo condition -- exactly the C1 row of the matrix
|
||||
- The verdict the class's Part III section states (Lecture 8): a \textbf{hexagonal modular monolith for the booking core} (MM and HX at $++$), EDA at the edges and PF for the batch runs as secondary, and microservices only when organisation size forces D11 to High -- the Monzo condition -- exactly the C1 row of the matrix
|
||||
|
||||
*Elements:*
|
||||
- two-column comparison (operational vs structural mitigation), deck-6 mirror-pair layout
|
||||
|
||||
Loading…
x
Reference in New Issue
Block a user