5-layer cake explained

The 5-Layer Cake Explained: NVIDIA’s infrastructure model for the EU market

Most AI infrastructure decisions in Europe fail somewhere other than where they were made. An accelerator is specified, then the site turns out to be unable to supply the rack density it needs.

A model is selected, then the logging its classification requires turns out to be unavailable from the layer underneath it.

Jensen Huang’s five-layer reading of AI infrastructure is a practical way to see those dependencies before committing to them. Energy sits at the base, then chips, infrastructure, models and applications, with every application pulling on each layer beneath it down to the power supply.

What the reference version leaves out is that in the EU each layer carries a constraint of its own.

Energy: what the EU makes you disclose

Every other layer can be traded against another. Energy cannot, because it sets the absolute ceiling on what a site is able to run.

In the EU that ceiling comes with a paper trail. Article 12 of the recast Energy Efficiency Directive, Directive (EU) 2023/1791, requires operators of data centres with an installed IT power demand of at least 500 kW to make energy performance information public.

Delegated Regulation (EU) 2024/1364 sets out what those operators report to the European database on data centres and how each figure is calculated, covering power usage effectiveness, water usage effectiveness, energy reuse factor and renewable energy factor.

Reports covering the preceding calendar year are due by 15 May, and the Commission has consulted on a second phase that converts the collected data into a rating and labelling scheme.

The practical effect is that efficiency above the threshold is a matter of record, so a claim about an efficient site can be checked instead of accepted.

Ask which database entry corresponds to the hall your workload will occupy, and what the reported PUE was against a load profile comparable to the one you intend to add.

A site holding a comfortable ratio on light mixed load will not hold it once dense accelerator racks arrive.

Chips: the constraints that bracket the decision

The layered reading makes clear that a chip decision is constrained from both directions at once. Below it, the power and cooling envelope of the site decides which rack densities are deployable at all.

Above it, the model you intend to run sets memory capacity and interconnect requirements that peak throughput figures do not capture.

That bracketing has a procurement consequence. Specifying a performance envelope travels better than specifying a part number, because the envelope can be tested against both the site below and the workload above, while a part number can only be matched or missed.

Specific accelerators are compared in the GPU selection post.

Infrastructure: where European options differ

Between chips and models sits everything that turns processors into usable capacity: power delivery, cooling, networking, storage and the orchestration that makes a cluster behave as a single machine. Huang’s term for the result is an AI factory.

Europe is building this layer as public infrastructure alongside the private market. The EuroHPC Joint Undertaking has selected 19 sites to host AI Factories, each acting as a national access point to AI-optimised HPC capacity for startups, SMEs and researchers.

The Lithuanian site, LitAI Factory, is led by Vilnius University and hosted at the LRTC VDC3 facility in Vilnius, with priority sectors named as cybersecurity, green energy, smart industry and digital health.

That changes the shape of the decision rather than simply adding capacity to the market. For research and early-stage development, subsidised national capacity is often the right first call.

For production workloads carrying residency or isolation commitments, allocation-based access rarely fits, and the choice moves to private or managed capacity. Where that line falls is worked through on the data sovereignty pillar.

Models: where the AI Act attaches

Model selection is where EU obligations land. The AI Act imposes requirements covering risk management, technical documentation, data governance, transparency and logging, varying by how a system is classified and whether an organisation acts as provider or deployer.

Those obligations resolve downward into infrastructure capabilities, which is why the choice cannot be made on benchmark performance alone. Logging requirements mean retained, queryable records of inference activity, held for a defined period and producible on request.

Technical documentation means traceability from a deployed artefact back to the dataset version and training run that produced it. Data governance means knowing where weights and training data physically sit, which constrains which layer-three options remain open.

None of that is achievable by adding a compliance step after deployment. It has to exist as instrumentation in the layers beneath, which is the argument for treating model operations as infrastructure rather than as a downstream activity.

The MLOps and AIOps post covers what that instrumentation involves in practice. A model requiring records your infrastructure layer cannot produce is not deployable on any accelerator.

Applications: the layer that should set the specification

At the top is where economic value is created, and requirements ought to flow downward from here rather than upward from whatever hardware was bought first.

  • Latency-bound applications such as real-time assistants set a response-time budget that constrains serving architecture, batching strategy and the physical distance between inference and user.
  • Throughput-bound batch work trades latency for cost per unit of output, which usually justifies different accelerators and a different scheduling approach entirely.
  • Applications handling regulated data import their controls into every layer beneath, from tenant isolation through to where model weights are stored.
  • Applications with bursty or seasonal demand sharpen the buy-versus-rent question, because owned capacity sized for peak sits idle for most of the year.

Working downward from here is what prevents capacity from being committed before the requirement justifying it has been written down. Who owns that translation in practice is covered in the AI infrastructure engineer post.

The value of reading AI as a stack is that it exposes the binding constraint rather than the interesting one, and in European deployments the binding constraint is rarely the model. The complete stack view is set out on the AI-ready infrastructure pillar.

Neurotechnology Cloud designs and operates these layers as one system through its AI Factory engineering service, from power and cooling envelope through to model deployment on dedicated European capacity.

If you are working out where the constraint sits in your own deployment, get in touch.

Share: 

Contact us

Interested in our products, custom solutions, or partnership opportunities? Have questions about our technologies or need more information before purchasing? Fill out the form, and our team will get back to you as soon as possible.