Auto-commit 2026-09-07 15:19: 9 files changed, 82 insertions(+), 96 deletions(-)

This commit is contained in:
herzogflorian 2026-09-07 15:19:17 +02:00
parent 161b95c757
commit 324633e30a
9 changed files with 82 additions and 96 deletions

View File

@ -172,7 +172,7 @@
\begin{itemize}\setlength\itemsep{4pt}
\item Not a rhetorical flourish but a \textbf{promissory note falling due}: Part I issued it as \textbf{Assumption A6} -- \emph{AI components extend the quality attribute space but do not change the method}. The bet in two sentences: everything AI does to software engineering can be absorbed by the apparatus you now own
\item If AI-bearing systems required a genuinely different method, the bet would be lost -- this part is where the claim must \textbf{survive contact with the evidence}
\item Roadmap -- \textbf{cases first, generalisation after}: two contradictory randomised experiments open Axis A (\S41); one concrete LLM call, wired wrongly and then rightly, opens Axis B (\S42); the matrix reading (\S43) and the emergent pattern (\S44) follow next week
\item Roadmap -- \textbf{cases first, generalisation after}: two contradictory randomised experiments open Axis A (\S 41); one concrete LLM call, wired wrongly and then rightly, opens Axis B (\S 42); the matrix reading (\S 43) and the emergent pattern (\S 44) follow next week
\end{itemize}
\end{frame}
@ -257,12 +257,12 @@
\end{frame}
\begin{frame}{The full empirical record, 2023--2025}
\scriptsize The two cases are the extreme corners of a larger record -- seven strands, 2023--2025, from RCTs to organisational telemetry and longitudinal code analysis; read every row \emph{setting first, finding second}:
\scriptsize The two cases are the extreme corners of a seven-strand record, 2023--2025 -- read every row \emph{setting first, finding second}:
\vspace{-0.15cm}
\renewcommand{\arraystretch}{0.78}%
\renewcommand{\arraystretch}{0.9}%
\begin{center}
\begin{tabular}{@{}>{\raggedright\arraybackslash}p{2.4cm}>{\raggedright\arraybackslash}p{3.8cm}>{\raggedright\arraybackslash}p{6.8cm}@{}}
\begin{tabular}{@{}>{\raggedright\arraybackslash}p{2.4cm}>{\raggedright\arraybackslash}p{3.4cm}>{\raggedright\arraybackslash}p{7.3cm}@{}}
\toprule
\textbf{Evidence} & \textbf{Setting} & \textbf{Finding} \\
\midrule
@ -277,30 +277,15 @@ Stack Overflow survey & $>$49{,}000 developers & 84\,\% use or plan to use AI; \
\end{tabular}
\end{center}
\vspace{-0.22cm}
\scriptsize \textcolor{codegray}{Caution (GitHub's own telemetry-plus-survey study): the best predictor of \emph{perceived} productivity is the suggestion acceptance rate, not the persistence of accepted code -- perception, not verified output; Case 2's perception gap is the controlled-trial demonstration of the same fact.}
\vspace{-0.1cm}
\scriptsize \textcolor{codegray}{Caution (GitHub's telemetry-plus-survey study): the best predictor of \emph{perceived} productivity is the suggestion acceptance rate, not the persistence of accepted code; Case 2's perception gap is the controlled-trial demonstration of this fact.}
\end{frame}
\begin{frame}{The system level: DORA 2024 and 2025}
\footnotesize
\begin{itemize}\setlength\itemsep{1pt}
\begin{itemize}\setlength\itemsep{3pt}
\item DORA measures neither task times nor perceptions but \textbf{delivery performance at the level of the organisation} -- throughput and stability -- exactly the level at which architecture acts
\item \textbf{2024} ($\sim$3{,}000 respondents; 75.9\,\% use AI for at least part of their work, roughly three quarters report productivity gains) -- a 25\,\% increase in AI adoption is associated with:
\begin{center}
\scriptsize
\renewcommand{\arraystretch}{0.85}%
\begin{tabular}{@{}ll@{}}
\toprule
\textbf{Gains} & \textbf{Losses} \\
\midrule
$+7.5\,\%$ documentation quality & $-1.5\,\%$ delivery throughput \\
$+3.4\,\%$ code quality & $-7.2\,\%$ delivery stability \\
$+3.1\,\%$ review speed & \\
\bottomrule
\end{tabular}
\end{center}
\vspace{-0.1cm}
Proposed mechanism is classical: more code per change, and larger batch sizes have been a documented risk driver for years
\item \textbf{2024} ($\sim$3{,}000 respondents; 75.9\,\% use AI for at least part of their work, roughly three quarters report productivity gains) -- a 25\,\% increase in AI adoption is associated with \textbf{gains} of $+7.5\,\%$ documentation quality, $+3.4\,\%$ code quality, $+3.1\,\%$ review speed -- and \textbf{losses} of $-1.5\,\%$ delivery throughput and $-7.2\,\%$ delivery stability. Proposed mechanism is classical: more code per change, and larger batch sizes have been a documented risk driver for years
\item \textbf{2025} ($\sim$5{,}000 respondents): adoption near saturation (90\,\%, median about two hours of daily use); more than 80\,\% report productivity gains; 30\,\% still express little or no trust in AI-generated code; the throughput association has \textbf{turned positive} as tools and practices matured -- the negative association with delivery stability \textbf{persists}
\item Central metaphor: \textbf{AI is an amplifier} -- it magnifies the strengths of well-run organisations and the dysfunctions of badly run ones
\item \emph{Individual acceleration and system-level performance are different quantities, and only the second one pays salaries}
@ -318,7 +303,7 @@ Stack Overflow survey & $>$49{,}000 developers & 84\,\% use or plan to use AI; \
\end{frame}
\begin{frame}{Reconciling the divergence: five moderator variables}
\footnotesize \textbf{So which study is wrong? Neither} -- resolving the contradiction \emph{is} the lesson: different populations (task novices vs.\ domain experts in their own code), different codebases (greenfield vs.\ mature), different tasks (bounded vs.\ real issues) -- the results never actually compete; the resolution requires reading \emph{study designs}, not abstracts. Indexed by their moderator variables, Case 1 and Case 2 sit at opposite corners of a five-dimensional design space -- and every other row of the record finds its place in the same coordinates.
\footnotesize \textbf{So which study is wrong? Neither} -- resolving the contradiction \emph{is} the lesson: different populations (task novices vs.\ domain experts in their own code), different codebases (greenfield vs.\ mature), different tasks (bounded vs.\ real issues) -- the results never actually compete; the resolution requires reading \emph{study designs}, not abstracts. Case 1 and Case 2 sit at opposite corners of a five-dimensional design space -- and every other row of the record finds its place in the same coordinates.
\vspace{0.1cm}
\footnotesize
@ -338,7 +323,7 @@ Organisation & small batches, test automation, loose coupling & large batches, w
\end{center}
\vspace{0.05cm}
\footnotesize The same technology yields $+55.8\,\%$ and $-19\,\%$ because the two cases differ on \textbf{every one of the five rows}.
\footnotesize The same technology yields $+55.8\,\%$ and $-19\,\%$ because the two cases differ on \textbf{all five rows}.
\end{frame}
\begin{frame}{Discussion: which setting is yours?}
@ -376,8 +361,8 @@ Organisation & small batches, test automation, loose coupling & large batches, w
\footnotesize
\begin{enumerate}\setlength\itemsep{2pt}
\item \textbf{Architecture quality gates AI gains} -- DORA 2025's core finding: teams in loosely coupled architectures with fast feedback loops convert AI adoption into throughput; tightly coupled systems with slow processes do not -- the AI-era echo of loosely coupled architectures and teams as the strongest predictor of continuous delivery performance \textcolor{codegray}{(the coupling finding of Lecture 11)}. In the theory's vocabulary: \textbf{D7} (evolvability) and \textbf{D9} (testability and deployability) gain weight in \emph{every} requirements profile -- architecture--application fit acquires a second reading: fit to a \emph{mode of work} in which change volume rises by an order of magnitude
\item \textbf{Architecture documentation becomes a control interface} -- ADRs, repository convention files, and machine-readable rules are no longer passive records; agents execute them on every run (\S41.6)
\item \textbf{Fitness functions become the operating licence for agents} -- an agent iterating against a dense test suite and CI-enforced architecture rules is contained; without them, every agent change is unpriced risk (\S41.7)
\item \textbf{Architecture documentation becomes a control interface} -- ADRs, repository convention files, and machine-readable rules are no longer passive records; agents execute them on every run (\S 41.6)
\item \textbf{Fitness functions become the operating licence for agents} -- an agent iterating against a dense test suite and CI-enforced architecture rules is contained; without them, every agent change is unpriced risk (\S 41.7)
\end{enumerate}
\vspace{0.1cm}
@ -386,7 +371,6 @@ Organisation & small batches, test automation, loose coupling & large batches, w
\end{keypoint}
\end{frame}
% ============================================
% AXIS A -- CONTROL INTERFACE AND GUARDRAILS
% ============================================
@ -410,17 +394,17 @@ Organisation & small batches, test automation, loose coupling & large batches, w
\# Portfolio Intelligence Platform -- agent instructions\\
\textbf{\#\# Architecture (binding; see docs/adr/)}\\
- Modular monolith, module boundaries enforced by CI\\
~~(see fitness\_functions/boundaries\_test.py). Do not add\\
~~cross-module imports; use the module's public API.\\
\hspace*{2\fontcharwd\font`0}(see fitness\_functions/boundaries\_test.py). Do not add\\
\hspace*{2\fontcharwd\font`0}cross-module imports; use the module's public API.\\
- All LLM access goes through gateway/ -- never call a\\
~~provider SDK from domain code (ADR-011).\\
\hspace*{2\fontcharwd\font`0}provider SDK from domain code (ADR-011).\\
\textbf{\#\# Verification (run before proposing changes)}\\
- make test~~~~~~~~~~\# unit + module-boundary rules\\
- make evals~~~~~~~~~\# eval harness; required for any\\
~~~~~~~~~~~~~~~~~~~~~\# change under prompts/ or gateway/\\
\hspace*{21\fontcharwd\font`0}\# change under prompts/ or gateway/\\
\textbf{\#\# No-go zones}\\
- ledger/ : append-only audit journal. Propose changes\\
~~as an ADR draft instead of editing code.
\hspace*{2\fontcharwd\font`0}as an ADR draft instead of editing code.
\end{tcolorbox}
\vspace{0.05cm}
@ -429,7 +413,7 @@ Organisation & small batches, test automation, loose coupling & large batches, w
\begin{frame}{Guardrails as the precondition for safe agent use}
\footnotesize
\begin{itemize}\setlength\itemsep{4pt}
\begin{itemize}\setlength\itemsep{1pt}
\item If verification is the scarce resource, then everything that \emph{automates} verification multiplies the value of AI tooling -- and everything that leaves verification informal converts AI speed into instability
\item \textbf{Test suites are the operating licence}: against a dense, fast test suite an agent can iterate -- wrong code fails immediately and is repaired or discarded at machine speed; without that net every agent-generated change ships \emph{unpriced risk} (DORA's ``strong version control and test automation'' amplifier pair)
\item \textbf{Architectural fitness functions fence the structure}: an objective integrity assessment of an architectural characteristic is the machine-readable form of an architecture decision -- dependency rules, cycle checks, module-boundary verification as CI gates were good practice before AI; with agents in the loop they are the mechanism by which an architect constrains \emph{a collaborator who never attends design meetings}
@ -445,7 +429,7 @@ Organisation & small batches, test automation, loose coupling & large batches, w
\item \textbf{Claude Code} (Anthropic): agentic CLI tool, research preview February 2025, GA May 2025; repository-level anchor \texttt{CLAUDE.md}
\item \textbf{Cursor} (Anysphere): AI-first IDE with an agent mode; the dominant tool among the METR study's experts
\item \textbf{GitHub Copilot}: Copilot Workspace retired May 2025; its concepts live on in the asynchronous \emph{Copilot coding agent} (issues to pull requests, in CI) and the synchronous IDE agent mode
\item \textbf{Devin} (Cognition): ``first AI software engineer'' (2024); 13.86\,\% SWE-bench in March 2024 triggered the agent wave; acquired Windsurf July 2025 -- rapid market consolidation
\item \textbf{Devin} (Cognition): ``first AI software engineer'' (2024); its 13.86\,\% SWE-bench result in March 2024 helped trigger the agent wave; acquired Windsurf July 2025 -- rapid market consolidation
\end{itemize}
\vspace{0.1cm}
@ -458,7 +442,7 @@ Organisation & small batches, test automation, loose coupling & large batches, w
\item \textbf{Two open standards matter more than any product, because they are architectural}
\item \textbf{Model Context Protocol (MCP)} -- Anthropic, November 2024: JSON-RPC; servers expose tools, resources, prompts; adopted by OpenAI, Google DeepMind, Microsoft in 2025; December 2025 to the Agentic AI Foundation under the Linux Foundation; over 10{,}000 public servers
\item \textbf{\texttt{AGENTS.md}} for project-level instructions
\item Vendor SDKs extract the agent loop as a library -- \emph{the bridge to Axis B}: the same building blocks that run SDLC agents also run runtime agent workflows (\S44, next week)
\item Vendor SDKs extract the agent loop as a library -- \emph{the bridge to Axis B}: the same building blocks that run SDLC agents also run runtime agent workflows (\S 44, next week)
\end{itemize}
\vspace{0.15cm}
@ -480,30 +464,30 @@ Organisation & small batches, test automation, loose coupling & large batches, w
\end{hinweisbox}
\end{frame}
\begin{frame}{Risks and responsibility (1/2): security, automation bias, skill}
\begin{frame}{Risks and responsibility (1/2): security, automation bias}
\footnotesize
\begin{itemize}\setlength\itemsep{4pt}
\item \textbf{Security} -- the evidence predates the agent wave and gains relevance with volume: roughly \textbf{40\,\%} of 1{,}689 Copilot-generated programs (89 security-relevant scenarios) contained CWE top-25 vulnerabilities; a user study: participants with an AI assistant wrote \textbf{less secure code while believing it more secure}; package hallucination (``slopsquatting''): across roughly 576{,}000 generations, \textbf{about a fifth} of recommended package references did not exist -- names an attacker can register pre-emptively. Consequence: SAST, dependency and secret scanning, licence checks in CI are not optional; \emph{security review capacity must scale with generation volume}
\item \textbf{Security} -- the evidence predates the agent wave and gains relevance with volume: roughly \textbf{40\,\%} of 1{,}689 Copilot-generated programs (89 security-relevant scenarios) contained CWE top-25 vulnerabilities; a user study: participants with an AI assistant wrote \textbf{less secure code on most tasks while believing it more secure}; package hallucination (``slopsquatting''): across roughly 576{,}000 generations, \textbf{about a fifth} of recommended package references did not exist -- names an attacker can register pre-emptively. Consequence: SAST, dependency and secret scanning, licence checks in CI are not optional; \emph{security review capacity must scale with generation volume}
\item \textbf{Automation bias}: over-trust in automated systems is a decades-old human-factors finding -- Perry et al.'s participants overestimated their security, METR's experts overestimated their speed
\item \textbf{Skill formation}: a randomised study of engineers learning a new library -- AI assistance reduced comprehension-test scores by roughly \textbf{17\,\%}; the usage pattern is the decisive moderator (conceptual questions preserved learning, wholesale delegation destroyed it); \textbf{entry-level developer positions are measurably declining}. For this module: the role being trained is the \emph{specifier, verifier, and architect} -- rebuild the competence ladder deliberately, including AI-free practice of fundamentals
\end{itemize}
\end{frame}
\begin{frame}{Risks and responsibility (2/2): accountability}
\small
\begin{itemize}\setlength\itemsep{5pt}
\item Legally and professionally, \textbf{the person who merges code answers for it}, regardless of what generated it; AI tools are not liability-bearing entities -- treat AI output as \emph{the contribution of an unknown third party}: mandatory review, provenance labelling, an explicit policy for permitted uses (DORA 2025: a clearly communicated AI policy \emph{first} among the seven amplifier capabilities)
\item AI may \emph{draft} an ADR; a nameable person decides, signs, and defends it \textcolor{codegray}{\footnotesize (deck 3)}
\begin{frame}{Risks and responsibility (2/2): skill, accountability}
\footnotesize
\begin{itemize}\setlength\itemsep{3pt}
\item \textbf{Skill formation}: a randomised study of engineers learning a new library -- AI assistance reduced comprehension-test scores by roughly \textbf{17\,\%}; the usage pattern is the decisive moderator (conceptual questions preserved learning, wholesale delegation destroyed it); \textbf{entry-level developer positions are measurably declining}. For this module: the role being trained is the \emph{specifier, verifier, and architect} -- rebuild the competence ladder deliberately, including AI-free practice of fundamentals
\item \textbf{Accountability}: legally and professionally, \textbf{the person who merges code answers for it}, regardless of what generated it; AI tools are not liability-bearing entities -- treat AI output as \emph{the contribution of an unknown third party}: mandatory review, provenance labelling, an explicit policy for permitted uses (DORA 2025: a clearly communicated AI policy \emph{first} among the seven amplifier capabilities)
\item AI may \emph{draft} an ADR; a nameable person decides, signs, and defends it \textcolor{codegray}{\scriptsize (deck 3)}
\end{itemize}
\vspace{0.1cm}
\vspace{0.05cm}
\begin{center}
\emph{\textcolor{bankblue}{\textbf{Architecture is an accountability performance, not a text-production performance.}}}
\small\emph{\textcolor{bankblue}{\textbf{Architecture is an accountability performance, not a text-production performance.}}}
\end{center}
\vspace{0.1cm}
\small
\begin{itemize}\setlength\itemsep{5pt}
\vspace{0.05cm}
\footnotesize
\begin{itemize}\setlength\itemsep{3pt}
\item \textbf{IP risk open but manageable}: \emph{Doe v.\ GitHub} -- the DMCA claim dismissed in 2024, licence-related claims continue; response: provider duplication filters and indemnification, licence scanning in CI, a documented residual risk in the governance record
\end{itemize}
\end{frame}
@ -514,7 +498,7 @@ Organisation & small batches, test automation, loose coupling & large batches, w
\begin{enumerate}\setlength\itemsep{2pt}
\item[(i)] the repository carries an \texttt{AGENTS.md}\,/\,\texttt{CLAUDE.md} in the spirit of the listing -- and you are expected to \textbf{keep it as current as code}
\item[(ii)] every architecture decision is an \textbf{ADR} -- agents may draft, but a named team member signs
\item[(iii)] agent-generated changes enter the main branch \textbf{only through the CI gate}: module-boundary fitness functions, the test suite, and -- for anything touching prompts or the gateway -- the \textbf{eval harness} of \S42.5
\item[(iii)] agent-generated changes enter the main branch \textbf{only through the CI gate}: module-boundary fitness functions, the test suite, and -- for anything touching prompts or the gateway -- the \textbf{eval harness} of \S 42.5
\item[(iv)] your project handbook contains a \textbf{one-page AI policy}: permitted tools, provenance labelling, review rules
\end{enumerate}
\vspace{0.05cm}
@ -522,7 +506,6 @@ Organisation & small batches, test automation, loose coupling & large batches, w
\end{projektbox}
\end{frame}
% ============================================
% AXIS B -- COMPONENT TYPES
% ============================================
@ -531,8 +514,8 @@ Organisation & small batches, test automation, loose coupling & large batches, w
\begin{frame}{Axis B opens: the news-sentiment call, wired the obvious way}
\emph{\textcolor{bankblue}{One of the platform's features is a single LLM call -- news in, sentiment out. Why not call it like any other function?}}
\vspace{0.15cm}
\small
\vspace{0.05cm}
\footnotesize
\begin{itemize}\setlength\itemsep{3pt}
\item Axis B moves AI \textbf{from the workshop into the product}; as always the case precedes the taxonomy: walk one concrete call end to end, watch what breaks, and name every break with a dimension you already own
\item \textbf{The feature}: when a user opens a portfolio, the platform fetches the latest news items for its positions and asks an LLM, per item -- \emph{is this news positive, negative, or neutral for this holding, and why?} One prompt, one structured answer: the simplest runtime AI component the course project owns
@ -542,7 +525,7 @@ Organisation & small batches, test automation, loose coupling & large batches, w
\vspace{0.1cm}
\begin{center}
\begin{tikzpicture}[
sysbox/.style={rectangle, draw, rounded corners=4pt, align=center, font=\scriptsize\sffamily, line width=0.8pt, fill=gray!15, draw=gray!60!black, minimum width=2.6cm, minimum height=0.7cm},
sysbox/.style={rectangle, draw, rounded corners=4pt, align=center, font=\scriptsize\sffamily, line width=0.8pt, fill=gray!15, draw=gray!60!black, minimum width=2.6cm, minimum height=0.6cm},
arr/.style={-{Stealth[length=2.5mm]}, thick, gray!60!black}
]
\node[sysbox] (h) at (0,0) {request handler};
@ -577,7 +560,7 @@ Organisation & small batches, test automation, loose coupling & large batches, w
\begin{enumerate}\setlength\itemsep{4pt}
\setcounter{enumi}{3}
\item \textbf{Drift (D7)}: the provider ships a new model version or deprecates the old one -- GA models carry deprecation windows of the order of \emph{six months} -- and the component's behaviour changes \textbf{without any local action}: no commit, no deployment, no reviewable diff. The feature's behaviour is now co-owned by a third party
\item \textbf{Injection (D6)}: the news article is untrusted input read by a component that cannot reliably separate instructions from data -- a crafted ``article'' can carry instructions to the model; the feature has quietly opened an attack surface that no classical threat model in the platform covers \textcolor{codegray}{(the attack surface in depth: \S42.6, next week)}
\item \textbf{Injection (D6)}: the news article is untrusted input read by a component that cannot reliably separate instructions from data -- a crafted ``article'' can carry instructions to the model; the feature has quietly opened an attack surface that no classical threat model in the platform covers \textcolor{codegray}{(the attack surface in depth: \S 42.6, next week)}
\end{enumerate}
\vspace{0.15cm}
@ -597,16 +580,16 @@ Organisation & small batches, test automation, loose coupling & large batches, w
\item Industry speaks of \textbf{compound AI systems} \textcolor{codegray}{(deck 6, C10)} for this reason: state-of-the-art results come from systems of models, retrievers, validators, and deterministic services, not a single model call
\end{itemize}
\begin{definitionbox}[AI runtime component]
\footnotesize A component of the delivered system whose output is produced by a \emph{learned or search-based model} rather than by explicitly programmed logic. Three types with systematically different profiles: \textbf{(a)} LLM components for analysis, extraction, and generation over unstructured input; \textbf{(b)} classical ML components for classification and regression; \textbf{(c)} optimisation components (LP/MIP and constraint solvers, metaheuristics). They differ exactly on the dimensions this theory measures and therefore demand \textbf{different integration forms}.
\footnotesize A component of the delivered system whose output is produced by a \emph{learned or search-based model} rather than by explicitly programmed logic. Three types with systematically different profiles: \textbf{(a)} LLM components for analysis, extraction, and generation over unstructured input; \textbf{(b)} classical ML components for classification and regression; \textbf{(c)} optimisation components (LP/MIP and constraint solvers, metaheuristics). The types differ exactly on the dimensions this theory measures -- determinism, latency, cost model, dominant risk, explainability -- and therefore demand \textbf{different integration forms}.
\end{definitionbox}
\end{frame}
\begin{frame}{The three AI component types and their profiles}
\scriptsize
\renewcommand{\arraystretch}{0.85}%
\renewcommand{\arraystretch}{0.95}%
\vspace{-0.2cm}
\begin{center}
\begin{tabular}{@{}p{2.0cm}p{3.5cm}p{3.3cm}p{3.5cm}@{}}
\begin{tabular}{@{}>{\raggedright\arraybackslash}p{2.0cm}>{\raggedright\arraybackslash}p{3.5cm}>{\raggedright\arraybackslash}p{3.5cm}>{\raggedright\arraybackslash}p{3.6cm}@{}}
\toprule
\textbf{Dimension} & \textbf{(a) LLM analysis / generation} & \textbf{(b) ML classification / regression} & \textbf{(c) Optimisation (LP/MIP/CP)} \\
\midrule
@ -626,23 +609,29 @@ Integration form & gateway $+$ async $+$ cache & serving endpoint $+$ MLOps pipe
\begin{frame}{Type (a): LLM components -- RAG, prompts, structured outputs}
\footnotesize
\begin{itemize}\setlength\itemsep{3pt}
\begin{itemize}\setlength\itemsep{1pt}
\item LLM components turn unstructured input -- documents, e-mails, reports -- into analyses, extractions, or generated text; \textbf{three engineering building blocks} define the type
\item \textbf{Retrieval-augmented generation (RAG)}: knowledge is moved out of the model weights into a swappable, versionable, inspectable \emph{data component} -- updated by re-indexing rather than retraining, with provenance through citable sources
\item RAG is an \emph{engineering} problem, not a model problem: case-study evidence documents seven recurring failure points (missing content, failed ranking of the relevant documents, extraction and formatting errors, incomplete answers) -- with the sobering observation that RAG robustness \emph{evolves} in operation rather than being designed in
\item RAG is an \emph{engineering} problem, not a model problem: case-study evidence documents seven recurring failure points (missing content, failed ranking, extraction and formatting errors, incomplete answers) -- with the sobering observation that RAG robustness \emph{evolves} in operation rather than being designed in
\item \textbf{Prompts are configuration artefacts}: version-controlled, regression-tested, behaviour-determining like code -- exactly the configuration-debt territory Sculley et al.\ mapped
\item \textbf{Structured outputs}: since 2024 provider APIs can enforce, via constrained decoding, that outputs conform to a developer-supplied JSON schema -- \emph{syntactic} correctness guaranteed; \emph{semantic} correctness remains to be verified (reference architecture, eval harness)
\item \textbf{Lifecycle risk is the provider}: GA models carry deprecation windows of the order of six months, shorter windows observed -- a hard-coded model name is a \emph{ticking dependency}: an architectural statement, not an operational one
\end{itemize}
\end{frame}
\begin{frame}{Types (b) and (c) -- perishable models, heavy solvers}
\footnotesize
\begin{itemize}\setlength\itemsep{1pt}
\begin{frame}{Type (b): classical ML components -- perishable models}
\small
\begin{itemize}\setlength\itemsep{4pt}
\item \textbf{(b)} Self-trained models (scoring, churn, fraud, forecasting) bring the full \textbf{nine-stage workflow} -- model requirements and data collection through training, evaluation, deployment, monitoring -- with dense feedback loops; characteristic problems: \emph{training/serving skew} (divergent data preparation, one of the most frequent production failure sources) and \emph{data/concept drift} (sudden, gradual, incremental, recurring)
\item \textbf{(b)} \textbf{A deployed model is a perishable good} -- monitoring and retraining are operating requirements, not options; tooling: feature stores with consistent online/offline views, model registries versioning model, data, code, and configuration together, the MLOps discipline \textcolor{codegray}{(maturity ladder: \S43, next week)}
\item \textbf{(c)} Routinely overlooked in the SE4AI literature but belongs in every advisory platform: LP/MIP solvers, constraint programming (CP-SAT dominated recent MiniZinc Challenges, a complete gold-medal sweep in 2024), stochastic metaheuristics -- the profile \emph{inverts} the LLM's: \textbf{deterministic but heavy}. Reproducible at fixed seed, thread count, time limit (run-to-run variability in practice); runtimes seconds to hours, often anytime behaviour $\to$ \textbf{asynchronous integration}: job queue, status polling, callback; never a synchronous call in a web request path
\item \textbf{(c)} Compensating strength: \textbf{provable explainability} -- optimality gap, dual values and shadow prices, and on infeasibility an irreducible infeasible subset (IIS): a minimal set of contradictory constraints as the explanation. In regulated domains the load-bearing argument for the project's division of labour: \emph{hard, auditable decisions belong to the solver and the deterministic services, not to the LLM}
\item \textbf{(b)} \textbf{A deployed model is a perishable good} -- monitoring and retraining are operating requirements, not options; tooling: feature stores (consistent online/offline views), model registries (model, data, code, configuration versioned together), the MLOps discipline \textcolor{codegray}{(maturity ladder: \S 43, next week)}
\end{itemize}
\end{frame}
\begin{frame}{Type (c): optimisation components -- heavy solvers}
\small
\begin{itemize}\setlength\itemsep{4pt}
\item \textbf{(c)} Routinely overlooked in the SE4AI literature but belongs in every advisory platform: LP/MIP solvers, constraint programming (CP-SAT: recent MiniZinc Challenges dominated, gold-medal sweep 2024), stochastic metaheuristics -- the profile \emph{inverts} the LLM's: \textbf{deterministic but heavy}. Reproducible at fixed seed, thread count, time limit (run-to-run variability in practice); runtimes seconds to hours, often anytime behaviour $\to$ \textbf{asynchronous integration}: job queue, status polling, callback -- never a synchronous call in a web request path
\item \textbf{(c)} Compensating strength: \textbf{provable explainability} -- optimality gap, dual values and shadow prices, on infeasibility an irreducible infeasible subset (IIS) as the explanation: a minimal set of contradictory constraints. In regulated domains the load-bearing argument for the project's division of labour: \emph{hard, auditable decisions belong to the solver and the deterministic services, not to the LLM}
\end{itemize}
\end{frame}
@ -658,7 +647,6 @@ Integration form & gateway $+$ async $+$ cache & serving endpoint $+$ MLOps pipe
\scriptsize \textcolor{codegray}{\textbf{Project transfer:} the sentiment call and the AdvisorAgent's insights are type (a); the Optimization service is type (c); type (b) has no instance in the project.}
\end{frame}
% ============================================
% AXIS B -- CONTAINMENT
% ============================================
@ -702,7 +690,7 @@ Integration form & gateway $+$ async $+$ cache & serving endpoint $+$ MLOps pipe
\begin{tikzpicture}[
sysbox/.style={rectangle, draw, rounded corners=4pt, minimum width=3.0cm, minimum height=0.9cm, align=center, font=\small\sffamily, line width=0.8pt},
core/.style={sysbox, fill=bankblue!20, draw=bankblue, font=\small\sffamily\bfseries, minimum height=2.4cm, minimum width=3.2cm},
gwpart/.style={sysbox, fill=aiviolet!15, draw=aiviolet, minimum width=3.4cm, minimum height=0.7cm, font=\scriptsize\sffamily},
gwpart/.style={sysbox, fill=aiviolet!15, draw=aiviolet, text width=3.3cm, minimum width=3.4cm, minimum height=0.75cm, font=\scriptsize\sffamily},
comp/.style={sysbox, fill=bankgreen!15, draw=bankgreen},
extern/.style={sysbox, fill=gray!15, draw=gray!60!black, minimum width=2.6cm},
guard/.style={sysbox, fill=bankred!10, draw=bankred},
@ -712,11 +700,11 @@ Integration form & gateway $+$ async $+$ cache & serving endpoint $+$ MLOps pipe
% Domain core (left)
\node[core] (core) at (0.4,1.7) {Deterministic\\domain core\\[2pt]{\scriptsize\mdseries decides and books;}\\{\scriptsize\mdseries no provider SDK imports}};
% Gateway internals (2+3 grid)
\node[gwpart] (assemble) at (7.0,2.6) {prompt assembly + schema validation};
\node[gwpart] (router) at (10.6,2.6) {model router (cheap $\rightarrow$ expensive cascade)};
\node[gwpart] (assemble) at (7.0,2.7) {prompt assembly + schema validation};
\node[gwpart] (router) at (10.9,2.7) {model router (cheap $\rightarrow$ expensive cascade)};
\node[gwpart] (cache) at (7.0,1.7) {semantic cache};
\node[gwpart] (breaker) at (10.6,1.7) {timeouts, circuit breakers, fallback chains};
\node[gwpart, minimum width=5.0cm] (cost) at (8.8,0.8) {cost telemetry per request / feature / tenant};
\node[gwpart] (breaker) at (10.9,1.7) {timeouts, circuit breakers, fallback chains};
\node[gwpart, text width=5.6cm, minimum width=5.8cm] (cost) at (8.95,0.7) {cost telemetry per request / feature / tenant};
\begin{scope}[on background layer]
\node[draw=aiviolet, line width=1pt, rounded corners=5pt, fill=aiviolet!5,
fit=(assemble)(router)(cache)(breaker)(cost),
@ -729,7 +717,7 @@ Integration form & gateway $+$ async $+$ cache & serving endpoint $+$ MLOps pipe
% Bottom row
\node[comp, minimum width=3.0cm] (queue) at (0.0,-1.7) {Async job queue\\{\scriptsize batching, backpressure, retries}};
\node[comp, minimum width=2.4cm] (workers) at (3.9,-1.7) {Worker pool\\{\scriptsize bounded concurrency}};
\node[guard, minimum width=5.2cm] (guard) at (8.8,-1.7) {Ontology / schema guard\\{\scriptsize entity resolution, domain axioms, citation check}};
\node[guard, minimum width=5.2cm] (guard) at (8.95,-1.7) {Ontology / schema guard\\{\scriptsize entity resolution, domain axioms, citation check}};
\node[evalb, minimum width=3.4cm] (eval) at (14.6,-1.7) {Eval harness\\{\scriptsize CI gate: prompts, models, providers}};
% Arrows
\draw[arr] ([yshift=0.5cm]core.east) -- node[above, font=\scriptsize\sffamily]{typed port} ([yshift=0.5cm]core.east -| gw.west);
@ -740,7 +728,7 @@ Integration form & gateway $+$ async $+$ cache & serving endpoint $+$ MLOps pipe
\draw[arr] (queue.east) -- (workers.west);
\draw[arr] (workers.north) |- ([yshift=-0.9cm]gw.west);
\draw[arr] (gw.south) -- node[right, font=\scriptsize\sffamily]{every output} (guard.north);
\draw[arr] (guard.south) -- (8.8,-2.75) -- (-2.3,-2.75) -- (-2.3,1.0) -- (core.west |- 0,1.0);
\draw[arr] (guard.south) -- (8.95,-2.75) -- (-2.3,-2.75) -- (-2.3,1.0) -- (core.west |- 0,1.0);
\node[font=\scriptsize\sffamily, anchor=north] at (3.3,-2.8) {validated result or rejection};
\draw[arr, dashed] (eval.west) -- (guard.east);
\draw[arr, dashed] (eval.north) |- ([yshift=-0.2cm]gw.south east);
@ -827,27 +815,27 @@ Integration form & gateway $+$ async $+$ cache & serving endpoint $+$ MLOps pipe
\begin{frame}{Example: an eval harness for the portfolio platform}
\scriptsize \textbf{Thresholds $=$ response measures}: mean F1 $\geq 0.92$ on the golden set $\cdot$ zero axiom violations $\cdot$ judge--human agreement $\kappa \geq 0.70$ (Cohen's chance-corrected measure) -- changing any of them is an architecture decision requiring an ADR; note what is \emph{absent}: no assertion demands an exact output string.
\vspace{0.05cm}
\begin{tcolorbox}[colback=gray!4!white, colframe=gray!55!black, boxrule=0.6pt, arc=2pt, top=3pt, bottom=3pt, left=6pt, right=6pt]
\vspace{0.02cm}
\begin{tcolorbox}[colback=gray!4!white, colframe=gray!55!black, boxrule=0.6pt, arc=2pt, top=1pt, bottom=1pt, left=6pt, right=6pt]
\scriptsize\ttfamily
GOLDEN = load\_cases("evals/portfolio\_extraction\_v3.jsonl")\\
def test\_extraction\_regression(gateway):\\
~~~~scores = [f1(gateway.extract(c.report), c.expected) for c in GOLDEN]\\
~~~~assert mean(scores) >= 0.92~~~\# statistical threshold\\
\hspace*{4\fontcharwd\font`0}scores = [f1(gateway.extract(c.report), c.expected) for c in GOLDEN]\\
\hspace*{4\fontcharwd\font`0}assert mean(scores) >= 0.92~~~\# statistical threshold\\
def test\_domain\_axioms(gateway, ontology):\\
~~~~answer = gateway.advise(sample\_portfolio())\\
~~~~for pos in answer.positions:\\
~~~~~~~~assert ontology.resolves(pos.isin), f"unknown: \{pos.isin\}"\\
~~~~total = sum(p.weight for p in answer.positions)\\
~~~~assert abs(total - 1.0) < 1e-6\\
~~~~for cit in answer.citations:\\
~~~~~~~~assert cit.passage in source\_text(cit.doc\_id)\\
\hspace*{4\fontcharwd\font`0}answer = gateway.advise(sample\_portfolio())\\
\hspace*{4\fontcharwd\font`0}for pos in answer.positions:\\
\hspace*{8\fontcharwd\font`0}assert ontology.resolves(pos.isin), f"unknown: \{pos.isin\}"\\
\hspace*{4\fontcharwd\font`0}total = sum(p.weight for p in answer.positions)\\
\hspace*{4\fontcharwd\font`0}assert abs(total - 1.0) < 1e-6\\
\hspace*{4\fontcharwd\font`0}for cit in answer.citations:\\
\hspace*{8\fontcharwd\font`0}assert cit.passage in source\_text(cit.doc\_id)\\
def test\_judge\_is\_calibrated(judge, human\_labels):\\
~~~~agreement = cohens\_kappa(judge.score(GOLDEN), human\_labels)\\
~~~~assert agreement >= 0.70~~~\# the judge's own eval
\hspace*{4\fontcharwd\font`0}agreement = cohens\_kappa(judge.score(GOLDEN), human\_labels)\\
\hspace*{4\fontcharwd\font`0}assert agreement >= 0.70~~~\# the judge's own eval
\end{tcolorbox}
\vspace{0.05cm}
\vspace{0.02cm}
\scriptsize \textcolor{codegray}{Runs in CI on every change to prompts, models, or the gateway, alongside the deterministic test suite.}
\end{frame}
@ -856,21 +844,20 @@ def test\_judge\_is\_calibrated(judge, human\_labels):\\
% ============================================
\section{Closing}
\begin{frame}{This week's exercise: AdvisorAgent + sub-agents behind the gateway}
\begin{frame}{This week's exercise: AdvisorAgent $+$ sub-agents behind the gateway}
\begin{projektbox}
\footnotesize \textbf{Coaching slot (1 lesson). Milestone M5} -- Multi-Agent Orchestration, Evaluation, and Hardening (weeks 12--13):
\begin{enumerate}\setlength\itemsep{2pt}
\item \textbf{Mandatory this week}: the AdvisorAgent orchestrates 2--3 sub-agents \emph{through contracts} -- every LLM call through the gateway port (ADR-011); no provider SDK import in domain code
\item \textbf{Ontology guard active on all insights}: entity resolution against the deterministic store, domain axioms, citation check -- unresolvable references are \emph{rejected}, not passed on
\item \textbf{LLM agents propose; deterministic services decide and book} -- keep the deterministic core free of LLM calls \emph{(the line that is graded)}
\item Axis A discipline as on the project-link slide: \texttt{AGENTS.md} current, ADRs signed, CI gate, AI policy
\end{enumerate}
\vspace{0.05cm}
\textbf{Week 13 turns the harness into a CI gate.}
\textbf{Axis A discipline as on the project-link slide; week 13 turns the harness into a CI gate.}
\end{projektbox}
\vspace{0.1cm}
\small The reference architecture of today is the topology you are wiring: typed port $\to$ gateway $\to$ queue $\to$ guard.
\scriptsize \textcolor{codegray}{The reference architecture of today is the topology you are wiring: typed port $\to$ gateway $\to$ queue $\to$ guard.}
\end{frame}
\begin{frame}{Summary}
@ -900,8 +887,8 @@ def test\_judge\_is\_calibrated(judge, human\_labels):\\
\begin{column}{0.42\textwidth}
\textcolor{bankblue}{\textbf{Reading}}
\begin{itemize}\small
\item this week: Part V, Sections 40--41, 42.1--42.5
\item ahead: Part V, Sections 42.6--42.7, 43--45
\item this week: Part V, sections 40--41, 42.1--42.5
\item ahead: Part V, sections 42.6--42.7, 43--45
\end{itemize}
\vspace{0.2cm}
@ -914,7 +901,6 @@ def test\_judge\_is\_calibrated(judge, human\_labels):\\
\end{columns}
\end{frame}
% ============================================
% END
% ============================================

View File

@ -344,7 +344,7 @@ Frame-by-frame plans for the remaining seven slide sets, derived strictly from t
- Right: \textbf{structural weakness} -- atomicity surrendered at the boundary (MS, D4): can only be \emph{made survivable}, under a condition most organisations do not meet
- The division of labour generalises: the veto rule does the heavy lifting, and the \textbf{``documented mitigation'' clause is where engineering knowledge -- not arithmetic -- enters the computation}
- The rest of the row follows the same mechanics (Section 33, Lecture 10); in particular \textbf{HX -- a delta discipline, not a competitor -- joins MM at $++$} by isolating the long-lived booking core from volatile channels and providers
- The verdict the class's Part III section states (Lecture 9): a \textbf{hexagonal modular monolith for the booking core} (MM and HX at $++$), EDA at the edges and PF for the batch runs as secondary, and microservices only when organisation size forces D11 to High -- the Monzo condition -- exactly the C1 row of the matrix
- The verdict the class's Part III section states (Lecture 8): a \textbf{hexagonal modular monolith for the booking core} (MM and HX at $++$), EDA at the edges and PF for the batch runs as secondary, and microservices only when organisation size forces D11 to High -- the Monzo condition -- exactly the C1 row of the matrix
*Elements:*
- two-column comparison (operational vs structural mitigation), deck-6 mirror-pair layout