Why the German Automotive Quality Model Does Not Work for Software
My first research project, back in 2002, was a collaboration with BMW. We were doing model-based testing of their MOST network master, using the approach Alex Pretschner had built in his PhD. We spent a lot of time with the BMW engineers building a detailed model in AutoFOCUS and then generating test cases from it. One of the things we found, and wrote up much later, was that a good share of the defects we found were found during modelling, not during test execution.
I have thought about that project on and off for more than twenty years. What strikes me now is something I did not notice at the time: everything around us in that building was world-class. The mechanical engineering was world-class. The process discipline was world-class. And the software was the part where nobody quite knew how to build it at the same world-class level.
Two decades later, that is still true. Launches slip because of software. Infotainment systems need a reboot. Over-the-air updates arrive late or not at all. Volkswagen has publicly downgraded CARIAD from developer to integrator and is buying its stack from Rivian and XPeng instead. Meanwhile, Manfred Broy's ICSE 2006 paper on the challenges in automotive software engineering and the roadmap paper by Pretschner, Broy, Krüger, and Stauner still read like current descriptions of the problem. A diagnosis that survives twenty years unchanged is either very good or very useless, and I think this one is very good. Which makes it worth asking why nothing moved.
The usual answers — they hired the wrong people, the culture is too slow, they were arrogant — are not wrong. But they are descriptions of what the problem looks like from outside, not explanations of what produces it. My hypothesis is more specific:
The German quality reputation is a manufacturing quality reputation. Its entire toolkit addresses variance in a reproduced process. Software has no reproduction variance, so every software defect is a design defect. The competence that produced the reputation is therefore aimed at a defect class that does not exist in the artefact, and the failure to transfer it is structural rather than a matter of effort or goodwill.
I do not want to claim I can prove this. What I want to do in this post is lay out five mechanisms that would make the hypothesis true, say for each of them which manufacturer seems to have avoided it, and — this is the part I actually care about — say what we would have to measure to know whether the mechanism is real. My honest assessment of our field is that we have an excellent descriptive literature on automotive software engineering and almost no causal one.
The procurement logic prices software as if it were a bracket. The supplier pyramid was a genuine advantage in mechanics, because mechanical interfaces are physical, visible, and stable. You specify a part, you tender it, you inspect it on arrival, you demand an annual price reduction. Every one of those steps misfires for software. Software has near-zero marginal cost, so a bill-of-materials model cannot price it, which pushes towards chronically undersized ECU hardware that the software then has to be contorted around. Suppliers deliver black-box binaries, so the OEM holds no source, no models, and no test artefacts, and therefore cannot do integration analysis or build a learning curve. And software couplings run through shared state, timing, and data — they are invisible at exactly the boundary where the contract is drawn.
Williamson's argument in The Economic Institutions of Capitalism is that the efficient governance form for a transaction depends on how specific the assets are and how cheaply performance can be verified. Where both are favourable, contracting against a specification wins. Where they are not, the transaction belongs inside the firm. Mechanical components sit in the first case; software sits in the second. Automotive procurement kept the governance form long after the properties that justified it had gone. The Center of Automotive Management makes the same link from the quality side, attributing rising recall volumes partly to the transfer of value creation to suppliers combined with platform strategies in which a late-discovered defect propagates across millions of vehicles.
Tesla's most consequential decision here was probably not centralised compute but source ownership: it treats controllers as white boxes it can rewrite, which is why it could re-flash firmware for alternative silicon during the 2021 chip shortage on a timescale that contracts made unavailable to everyone else. And notice what VW's move to Rivian concedes. Not that the software was too hard. That the contractual and organisational form could not produce it.
What would settle this? A study that holds product segment and variant count constant and varies the degree of in-house source control against downstream field quality. Recall databases are public, and software-attributable recalls can be coded. The hard part is operationalising "degree of source control" per vehicle programme and getting enough programmes to compare. I think this is the single most valuable empirical study nobody has done in our area.
The architecture is the org chart, and nobody wants to redraw the org chart. Conway's observation has become the mirroring hypothesis, with quantitative support from MacCormack, Baldwin and Rusnak. A vehicle E/E architecture with more than a hundred distributed ECUs is not primarily a technical choice. It is a picture of a domain- and supplier-partitioned organisation. Moving to zonal or centralised compute therefore means reorganising across departmental and contractual boundaries, which is far more expensive than any technical migration and threatens people who have real power.
Colfer and Baldwin use the car industry as their example of failed mirror-breaking: when firms push modularisation, nature sometimes stops them by exposing interdependencies nobody knew were there. Henderson and Clark's theory of architectural innovation fits almost too neatly — incumbents fail not for lack of component knowledge but when the architecture changes, because their information filters and communication channels encode the old architecture. Excellent components, structurally blocked architecture. Detlef Zerfowski framed the same tension as verticalisation versus horizontalisation in a talk I made sketchnotes of back in 2018, and I am not sure the industry has moved much since.
BMW's Neue Klasse is the most interesting test case here, because it is an incumbent attempting a genuine inverse Conway manoeuvre rather than a new entrant sidestepping the problem. Whether it works is not yet decidable. The Chinese entrants simply never accumulated the structure that has to be broken, which makes them a weak comparison: they are not better at mirror-breaking, they just had nothing to break.
This one is testable and essentially untested in our domain. The method exists: design structure matrices of the functional architecture compared against the organisational and contractual network, across several model generations of the same OEM. If anyone has that historical data and wants to open it up, I would take that collaboration tomorrow.
Our process standards are optimised for auditability, not for learning. Automotive SPICE, ISO 26262 and the V-model are built to demonstrate something to a third party. That is a legitimate goal, but it is a different goal from short feedback cycles, and the two trade off against each other. Because process maturity level is a contractual object between OEM and supplier, the rational response is to optimise for the assessment rather than for the product. Requirements inherit the same distortion. In the Lastenheft/Pflichtenheft model a requirement is a legal artefact: changing one means a variation order and a price negotiation, not a learning event. So the process penalises exactly the requirement volatility that our NaPiRE surveys keep finding as a dominant cause of project problems — penalises it without removing it. In recent work with colleagues on test case specification techniques and system testing tools actually in use in the automotive industry, the gap between what the standards imply and what teams do is again quite visible.
Toyota is the instructive comparison, because Toyota is also a process-heavy incumbent and did not respond by declaring itself agile. Woven by Toyota built Arene as a platform plus SDK plus tooling layer, which debuted in the 2026 RAV4 after more than five years of work. The bet is that you shorten feedback cycles by industrialising the development and validation infrastructure. That is a substantively different theory of the problem from CARIAD's, and it is running as a live experiment right now.
Here is the uncomfortable study: has anyone ever shown that Automotive SPICE capability level predicts field software quality? An enormous amount of industrial effort rests on the assumption that it does. Linking assessed capability levels to software-attributable recalls and warranty data, with honest controls for programme complexity, would be a service to the community even if — especially if — the answer is no.
Release cadence is locked to the vehicle programme, and part of that lock is real. Two forces do this. Internally, software rides on milestone gates designed for tooling lead times and ramp-up curves, with frozen baselines, which structurally guarantees late integration. Externally, ISO 26262, ISO/SAE 21434 and UNECE R155/R156 place a genuine cost on every release touching a safety-relevant function. This second force is the one most often waved away in commentary, and it is the most concretely real of everything in this post. The customer-visible consequence is measurable: a DAT survey found that 28 percent of vehicles under three years old had already had at least one defect in infotainment, connectivity or assistance systems.
Hyundai's Pleos is an attempt to decouple the software release train from the vehicle programme by making the platform the unit of planning. The Chinese OEMs get the same decoupling largely by having shorter vehicle programmes to begin with, which confounds the comparison badly.
Nobody has quantified the regulatory release tax. What is the marginal certification cost and calendar time of an OTA update as a function of the ASIL of the touched function, and how much of that is irreducible versus an artefact of how we currently produce homologation evidence? That is answerable with document analysis and practitioner interviews. It would also separate the "we are slow because of regulation" claim from the "we say regulation, we mean organisation" claim, and I would genuinely like to know which one is true.
You cannot absorb competence you cannot evaluate. Cohen and Levinthal's absorptive capacity argument says a firm can only take up external knowledge in proportion to related prior internal knowledge. After twenty-five years of outsourcing software, German OEMs lack precisely the internal base needed to evaluate, integrate and lead acquired software competence. This explains something otherwise puzzling: hiring thousands of developers did not work. Brooks' law is part of it, but the bigger part is a management layer socialised in mechanical engineering that assesses software progress through status reports rather than through running integrated builds. There is a status dimension too. Powertrain and chassis were the career paths; software was a supplier topic, and the reward structures encoded that long after the value in the product had moved.
The contrast I find most instructive is not Tesla but Toyota again. Woven was set up as a separate entity with its own culture and a bounded technical mandate. CARIAD was given an enormous scope and, by most accounts, insufficient authority relative to the brands. Same structural idea, very different outcome — which is exactly what makes the pair worth studying properly rather than anecdotally.
The Chalmers group around Eric Knauss and Patrizio Pelliccione has done the closest thing we have with their work on inter-organisational communication in automotive, largely with Volvo Cars. But it is descriptive. What is missing is work on managerial absorptive capacity specifically: what do OEM decision-makers actually use as evidence of software progress, how well do those signals track the real integration state, and does the gap predict schedule slip? That is tractable, and it would embarrass a lot of people, which is presumably why it does not exist.
Now the part where I argue against myself, because it would be far too easy to make this a story of incumbent stupidity. Three things make the German problem objectively harder than a smartphone stack. Andreas Vogelsang and Steffen Fuhrmann's analysis of a real vehicle functional architecture found that at least 85 percent of the analysed features depend on each other, and that developers were unaware of a large share of those dependencies when they were modelled only at architecture level. Platform kits times trim levels times markets times model years produce a configuration space that makes exhaustive integration testing combinatorially impossible — a software product line problem at a scale the pure-play EV makers have simply chosen not to have. And fifteen-plus years of field support and per-market type approval on software change is a constraint Tesla did not face for most of its existence.
Tesla and the Chinese entrants are not only better at software. They also opted out of part of the problem. Any causal claim in this space that does not control for variant count and programme longevity is worth very little, including mine.
So where does that leave us? If I had to pick the five studies I would most like to see funded, they would be:
- Software-attributable recall and warranty data against degree of in-house source control, controlling for segment and variant count.
- Longitudinal design structure matrices of functional architecture against organisational and contractual structure, across model generations at one OEM.
- Automotive SPICE capability level against field software quality, with honest controls.
- The marginal certification cost and calendar time of an OTA update by ASIL, separating irreducible cost from current evidence-production practice.
- What signals OEM leadership treats as evidence of software progress, and how well those signals track integration reality.
All five need industry data access that is difficult but not impossible. All five would produce findings that someone could act on the next morning. And all five are, I think, more useful than another survey confirming that automotive software development is hard.
I am aware that this post is closer to a set of hypotheses than to a set of results, and that is deliberate. I would be very happy if it started an argument. If you work at an OEM or a Tier 1 and one of these questions is either something you would like answered or something you are slightly afraid of the answer to, please get in touch — both are good reasons to talk.
References
- Apel, S., Batory, D., Kästner, C., & Saake, G. (2013). Feature-Oriented Software Product Lines. Springer.
- Baldwin, C. Y., & Clark, K. B. (2000). Design Rules, Vol. 1: The Power of Modularity. MIT Press.
- Brooks, F. P. (1987). No Silver Bullet: Essence and Accidents of Software Engineering. IEEE Computer, 20(4), 10–19.
- Broy, M. (2006). Challenges in Automotive Software Engineering. ICSE '06, 33–42.
- Broy, M., Krüger, I. H., Pretschner, A., & Salzmann, C. (2007). Engineering Automotive Software. Proceedings of the IEEE, 95(2), 356–373.
- Cohen, W. M., & Levinthal, D. A. (1990). Absorptive Capacity: A New Perspective on Learning and Innovation. Administrative Science Quarterly, 35(1), 128–152.
- Colfer, L. J., & Baldwin, C. Y. (2016). The Mirroring Hypothesis: Theory, Evidence, and Exceptions. Industrial and Corporate Change, 25(5), 709–738.
- Conway, M. E. (1968). How Do Committees Invent? Datamation, 14(5), 28–31.
- Ebert, C., & Favaro, J. (2017). Automotive Software. IEEE Software, 34(3), 33–39.
- Eliasson, U., Heldal, R., Knauss, E., & Pelliccione, P. (2015). The Need of Complementing Plan-Driven Requirements Engineering with Emerging Communication: Experiences from Volvo Car Group. RE '15.
- Henderson, R. M., & Clark, K. B. (1990). Architectural Innovation: The Reconfiguration of Existing Product Technologies and the Failure of Established Firms. Administrative Science Quarterly, 35(1), 9–30.
- Hohl, P., Münch, J., Schneider, K., & Stupperich, M. (2016). Forces that Prevent Agile Adoption in the Automotive Domain. PROFES 2016.
- Kasauli, R., Knauss, E., Horkoff, J., Liebel, G., & de Oliveira Neto, F. G. (2021). Requirements Engineering Challenges and Practices in Large-Scale Agile System Development. Journal of Systems and Software, 172.
- MacCormack, A., Baldwin, C., & Rusnak, J. (2012). Exploring the Duality Between Product and Organizational Architectures: A Test of the "Mirroring" Hypothesis. Research Policy, 41(8), 1309–1324.
- Méndez Fernández, D., Wagner, S., Kalinowski, M., Felderer, M., Mafra, P., Vetrò, A., et al. (2017). Naming the Pain in Requirements Engineering: Contemporary Problems, Causes, and Effects in Practice. Empirical Software Engineering, 22(5), 2298–2338.
- Pelliccione, P., Knauss, E., Heldal, R., Ågren, S. M., Mallozzi, P., Alminger, A., & Borgentun, D. (2017). Automotive Architecture Framework: The Experience of Volvo Cars. Journal of Systems Architecture, 77, 83–100.
- Pretschner, A., Broy, M., Krüger, I. H., & Stauner, T. (2007). Software Engineering for Automotive Systems: A Roadmap. FOSE '07, 55–71.
- Staron, M. (2021). Automotive Software Architectures: An Introduction (2nd ed.). Springer.
- Vogelsang, A., & Fuhrmann, S. (2013). Why Feature Dependencies Challenge the Requirements Engineering of Automotive Systems: An Empirical Study. RE '13.
- Williamson, O. E. (1985). The Economic Institutions of Capitalism. Free Press.
- Zyberaj, D., Hirmer, P., Aiello, M., & Wagner, S. (2026). Test Case Specification Techniques and System Testing Tools in the Automotive Industry. Journal of Systems and Software, 235.