Kostenloses Live-Webinar: VergabeHero in Aktion erleben.

Verfahrensbrief und Leistungsbeschreibung - Soofi_Compute_Tender_Specification_v1.8.pdf

GPU- und CPU-Rechenleistung für das Training eines großen Sprachmodells

Extrahierter Dokumenttext · Stand: 07.10.2026, 19:19 (Europe/Berlin)

Herkunft: www.deutsche-evergabe.de

Tabellen, Layout und Zeichen können bei der Extraktion abweichen. Maßgeblich ist die Originaldatei.

Originaldatei öffnen

[Seite 1]

Specification for the Procurement of GPU and CPU

Compute

for the Pretraining and Post-Training of Large

Language Models

Two lots: Lot 1 DGX/HGX B200/B300 GPU nodes · Lot 2 GB200/GB300 NVL72

racks

Gottfried Wilhelm Leibniz Universität Hannover (LUH), L3S Research Center Europe-wide call for tender

Page 1 of 28

[Seite 2]

Table of Contents

  1. General Information ........................................................................................................................... 3 1.1. Subject of the Tender (Summary) ................................................................................................ 3 1.2. Lots and Estimated Timeline ........................................................................................................ 4 1.3. Service Agreement ....................................................................................................................... 6 1.3.1. Place of Performance and Billing Address ............................................................................ 6 1.3.2. Period of Performance (Base Period) and Option ................................................................ 6 1.3.3. Contact Points ....................................................................................................................... 6 1.3.4. General Terms and Conditions for Bids and Contracts ......................................................... 6 1.3.5. Language ............................................................................................................................... 7 1.3.6. Prices ..................................................................................................................................... 7 1.3.7. Payment Terms ..................................................................................................................... 7 1.3.8. Jurisdiction ............................................................................................................................ 8 1.3.9. Conditionality ........................................................................................................................ 8 1.3.10. Data Center Location .......................................................................................................... 8 1.3.11. Bids per Lot ......................................................................................................................... 8 1.3.12. Variant bids (Nebenangebote) ............................................................................................ 8 1.4. Requirements ............................................................................................................................... 8
  2. Evaluation Criteria ............................................................................................................................. 22
  3. Annex ................................................................................................................................................ 23 3.1. Annex A — Technical Justification of the Hardware Requirements .......................................... 23 3.1.1. InfiniBand Interconnect ...................................................................................................... 23 3.1.2. NVIDIA Software Stack ........................................................................................................ 23 3.1.3. Blackwell Generation for Lot 1 and Exclusion of Hopper, Ampere and Older GPUs .......... 23 3.1.4. NVL72 Rack-Scale Systems for Lot 2 ................................................................................... 24 3.2. Annex B — Example of Price Evaluation (F1a + F1b), Lot 2 ....................................................... 25 3.3. Annex C — Data Protection and Location Requirements .......................................................... 26 3.4. Annex D — Supporting Evidence for Technical Requirements .................................................. 26 3.4.1. Compute-to-Storage Causal Chain ...................................................................................... 26 3.4.2. Reference List ...................................................................................................................... 27

Page 2 of 28

[Seite 3]

1. General Information

1.1. Subject of the Tender (Summary)

As part of a publicly funded European project, the Gottfried Wilhelm Leibniz Universität Hannover (LUH) publishes this Europe-wide call for tender for GPU and CPU compute for the pretraining and post-training of large language models (LLMs). The procurement is divided into two lots that correspond to two technically distinct system classes:

• Lot 1 — GPU nodes according to the NVIDIA DGX/HGX B200 or B300 reference architecture (8 GPU modules per node) or newer. Finishing and post-training of mixture-of-experts (MoE) models with more than 100 billion parameters, evaluation, synthetic data generation with generator models of about 400 billion parameters, distillation of smaller model variants.

• Lot 2 — Rack-scale systems according to the NVIDIA GB200 or GB300 NVL72 reference architecture or newer. Pretraining and post-training (e.g., reinforcement learning with verifiable rewards) of an MoE model with more than 350 billion parameters. The 72-GPU NVLink domain is a technical necessity for this model size (Annex A.4).

Bids may be submitted for one or both lots; each lot is evaluated and awarded separately (Section 1.3.11). The binding minimum compute requirement is 1.5 × 10²⁵ FLOPs for Lot 1 and 2.4 × 10²⁵ FLOPs for Lot 2. FLOPs are defined as (GPU-hours delivered) × (vendor-specified non-sparse peak BF16 Tensor Core throughput per GPU). The definition refers to peak-equivalent FLOPs as specified by the manufacturer, not to sustained FLOPs. The FLOP equivalent of the offered GPU model is calculated from its vendor-specified peak BF16 throughput; only the GPU models and system classes listed in A1 are admissible.

Key figureLot 1 — GPU nodes (DGX/HGXLot 2 — NVL72
B200/B300 class)
Reference GPUNVIDIA B200 in nodes of the DGX/HGX B200 reference architecture (2.25 × 10¹⁵ FLOPs/s BF16 dense per GPU); DGX/HGX B300 or newer NVIDIA reference architectures of the same class (e.g. DGX Rubin NVL8) acceptableNVIDIA GB200 NVL72 (2.5 × 10¹⁵ FLOPs/s BF16 dense per GPU); GB300 NVL72 or newer NVIDIA rack-scale reference architectures (e.g. Vera Rubin NVL144) acceptable
Binding minimum compute1.5 × 10²⁵ FLOPs ≈ 1.85 million B200 GPU-hours2.4 × 10²⁵ FLOPs ≈ 2.67 million GB200 GPU-hours (≈ 2,048 GPUs for at least 54 days)
Reference configuration1,024 GPUs (128 nodes) for the first 65 days, then 128 GPUs (16 nodes) to the end of the term (155 days) = 1,873,920 GPU-hours = 1.52 × 10²⁵ FLOPs2,048 GPUs (29 NVL72 racks, spares per A5) for 56 days from the start date (reference: 15 January to 11 March 2027) = 2,752,512 GPU-hours = 2.48 × 10²⁵ FLOPs

Page 3 of 28

[Seite 4]

Key figureLot 1 — GPU nodes (DGX/HGXLot 2 — NVL72
B200/B300 class)
Minimum concurrent GPUs (F)≥1,024 (128 nodes) during the first 65 days; ≥128 (16 nodes) thereafter2,048 (29 NVL72 racks) until the guaranteed FLOPs are delivered
Start / end of period of performanceStart date offered by the bidder, not before 1 December 2026 and not later than 15 January 2027; term 155 days (reference: 1 December 2026 to 4 May 2027), preceded by a two-week test phase (G2)Start date offered by the bidder, not before 15 January 2027 and not later than 1 April 2027; term until the guaranteed FLOPs are delivered (reference: 56 days, 15 January to 11 March 2027), preceded by a two-week test phase (G2)
OptionUp to three months after the base period, 128 GPUsUp to three months after the base period, 2,048 GPUs
Price ceiling16,400,000 € (incl. VAT of 19%)26,100,000 € (incl. VAT)
CPU≥15 million core-hours; partition of ≥4,096 physical corespartition of ≥8,192 physical cores (RL sandboxes)
Storage≥1 PB parallel filesystem plus ≥1 PB S3-compatible object storage≥1 PB parallel filesystem plus ≥1 PB object storage (checkpoints of up to about 8 TB each)
LocationEU/EEA or country with an EU adequacy decision (Annex C)EU/EEA or country with an EU adequacy decision (Annex C)

The allocation of the compute across tasks is at the discretion and risk of the contracting authority and does not concern the bidder. Lot 1 covers finishing and post-training of a model with more than 100 billion parameters, evaluation, synthetic data generation for European languages and domains, and distillation of smaller variants; Lot 2 covers the pretraining of a model with more than 350 billion parameters.

All quantitative requirements in Section 1.4 are derived from this workload and from measurements on the same model family; the derivation is given in Annex A and Annex D. The requirements are proportionate to the objective and do not exceed what is necessary to complete the programme within the period of performance.

1.2. Lots and Estimated Timeline

The two lots are independent contracts and may be awarded to different bidders. Their workloads are coupled only through data: training data for Lot 2 (about 150 TB) may be prepared in Lot 1 and are transferred before the start of Lot 2; checkpoints of the Lot 2 model (up to about 8 TB each) may be transferred from Lot 2 to Lot 1 for evaluation. Both lots must therefore provide data ingress and egress free of charge (C4). The timelines below are indicative; equivalent or superior timelines using other eligible GPUs are acceptable if they deliver the binding minimum FLOPs within the period of performance.

Page 4 of 28

[Seite 5]

Lot 1 (reference: NVIDIA B200, DGX/HGX nodes)

PeriodFLOPsGPUsGPU-hoursDescription
Test phase, two weeks before the start (reference 17 to 30 Nov 2026)–subset–Data ingest, validation runs on a subset of the nodes (G2, B6); not part of the term and not invoiced
Days 1–65 (reference 1 Dec 2026 to 3 Feb 2027)1.29 × 10²⁵1,0241,597,440Finishing and post-training of the >100B model; evaluation
Days 66–155 (reference 4 Feb to 4 May 2027)0.22 × 10²⁵≥128276,480Evaluation, synthetic data for domains, distillation of smaller variants
Total1.52 × 10²⁵1,873,920
Option (3 months after the base period)up to 0.23 × 10²⁵128up to 282,624Up to three months at the offered unit price (1.3.2); not part of the binding minimum

Lot 2 (reference: NVIDIA GB200 NVL72)

PeriodFLOPsGPUsGPU-hoursDescription
Test phase, two weeks before the start (reference 1 to 14 Jan 2027)–subset–Data ingest, validation runs on a subset of the nodes (G2, B6); not part of the term and not invoiced
56 days from start (reference 15 Jan – 11 Mar 2027)2.48 × 10²⁵2,0482,752,512Pretraining of the MoE model in NVFP4
Option (3 months after the base period)up to 4.07 × 10²⁵2,048up to 4,521,984Up to three months at the offered unit price (1.3.2); continued pretraining and post-training; not part of the binding minimum

Page 5 of 28

[Seite 6]

1.3. Service Agreement

1.3.1. Place of Performance and Billing Address Services are provided online. Billing address: Gottfried Wilhelm Leibniz Universität Hannover, Forschungszentrum L3S, Appelstr. 4, 30167 Hannover, Germany, attn. Prof. Dr. Wolfgang Nejdl.

1.3.2. Period of Performance (Base Period) and Option Lot 1: The bidder states the start date, not before 1 December 2026 and not later than 15 January 2027; the term runs 155 days from that date and is preceded by a test phase of at least two weeks (G2).

Lot 2: The bidder states the start date, not before 15 January 2027 and not later than 1 April 2027; the term runs until the guaranteed FLOPs have been delivered (reference: 56 days) and is preceded by a test phase of at least two weeks (G2). For both lots, the contracting authority may postpone the start by up to 30 days by notice at least four weeks before the offered date; the term then shifts accordingly. If the complete cluster has not been made available and accepted (G2) within 30 days after the start date, the contracting authority may terminate the contract for the lot; advance payments are repaid under the performance guarantee (1.3.7).

For each lot, the contracting authority may extend the agreement by up to three months following the base period (option), under identical conditions and at the offered unit price per GPU-hour of the base period (Lot 1: 128 GPUs, Lot 2: 2,048 GPUs). The option is included in the estimated contract value. Options may be exercised at any time after conclusion of the contract and at the latest four weeks before the end of the base period; option months are invoiced upon call-off (1.3.7). The binding minimum FLOPs and the evaluation (Section 2) refer to the base period. All services, including any optional services, shall be completed no later than 30 September 2027. The contracting authority may also exercise the option for periods shorter than a full month, including periods of individual days or weeks. In such cases, the price for the respective option period shall be calculated pro rata based on the applicable unit price per GPU-hour.

1.3.3. Contact Points All communication must be conducted exclusively via the electronic procurement platform www.deutsche-evergabe.de.

1.3.4. General Terms and Conditions for Bids and Contracts The following contractual documents shall apply to the procurement and shall form the basis for the contractual relationship between Leibniz Universität Hannover and the successful bidder:

• EVB-IT Cloud Contract;

• EVB-IT Cloud General Terms and Conditions (EVB-IT Cloud AGB);

• Additional Contractual Conditions (ZVB) of the State of Lower Saxony for the performance of supplies and services. (Zusätzliche Vertragsbedingungen (ZVB) des Landes Niedersachsen für die Ausführung von Lieferungen und Leistungen)

Page 6 of 28

[Seite 7]

The EVB-IT Cloud Contract will be finalised following the award decision with the bidder identified as the successful bidder. The final version of the contract shall be based on the procurement documents and the terms and conditions specified above.

By submitting a bid, bidders acknowledge and accept the applicability of the contractual documents specified above.

1.3.5. Language The bidder must prepare the offer, including all annexes and supporting documents, in German or English. All correspondence with the contracting authority, as well as the contract and negotiations, may be conducted in German or English.

1.3.6. Prices The price ceilings of 16,400,000 € (incl. VAT) for Lot 1 and 26,100,000 € (incl. VAT) for Lot 2 are determined by budgetary regulations and the contracting authority’s approved budget plan. They represent the binding maximum compensation for the tendered service of each lot; bids above the ceiling of a lot are excluded from that lot. A reduced VAT rate does not apply. The ceilings cover all services specified in this document, including CPU resources, storage, data transfer, support and login nodes (see 1.3.7).

The price weighting of 60% reflects the purpose of the procurement, to maximise delivered compute within a fixed budget; the qualitative minimum requirements are secured by the exclusion criteria (F).

1.3.7. Payment Terms Prices in the price sheet shall be quoted net (excl. VAT) if not stated otherwise; the price ceilings (1.3.6) and the price evaluation (F1b) refer to the grand total including VAT (19%).

The quoted prices shall cover the provision of fully operational cloud computing resources, including GPUs and CPUs, memory, storage, network capacity, general system requirements and all other requirements specified in this procurement specification, as well as all associated services, including provisioning, configuration and support.

The full contract price of the base period of each lot shall be invoiced no later than 1 December 2026, subject to the provision of a free-of-charge performance guarantee (Bankbürgschaft) covering 100% of the respective contract amount.

The performance guarantee is required to secure the proper fulfilment of the contractual obligations, including any claims of the contracting authority arising in connection with amounts paid in advance for services that have not yet been fully rendered or utilised. The performance guarantee shall remain effective at least until the services covered by the respective payment have been fully rendered or utilised and all relevant contractual claims of the contracting authority have been settled.

The performance guarantee shall be provided in the form of a bank guarantee / bank surety (Bankbürgschaft) issued by an authorised bank or financial institution. A corporate guarantee, parent company guarantee or other form of security shall not be considered equivalent to a bank guarantee.

Page 7 of 28

[Seite 8]

The bank guarantee shall be provided to the contracting authority in the original prior to payment of the relevant invoice. All costs associated with the provision and maintenance of the bank guarantee shall be borne by the contractor.

Option months shall be paid upon call-off; the performance guarantee shall also cover the amounts paid for option months.

1.3.8. Jurisdiction The jurisdiction is Hannover, Germany.

1.3.9. Conditionality The funds for this procurement have been granted to the contracting authority. According to the grant decision, they are released for payment by the funding agency upon submission of the selected offer; the contracting authority will apply for this release immediately after the evaluation of the offers, and the award will be made after the release. The exercise of the option (1.3.2) is conditional upon the availability of additional funds.

1.3.10. Data Center Location For legal and data protection reasons, all compute services must be provided exclusively within the EU/EEA or in a third country with an adequacy decision by the European Commission under GDPR. This requirement is mandated by European data protection law (GDPR, Schrems II); see Annex C.

1.3.11. Bids per Lot Bidders may bid for one lot or for both lots and must submit a separate price and a separate requirements sheet (Section 1.4) for each lot. Each lot is evaluated and awarded separately (Section 2). A bidder may state a combined discount that applies only if both lots are awarded to that bidder; the discount is taken into account in the price evaluation of both lots in proportion to the lot prices, provided that, after applying the combined discount, the bidder’s offer achieves the highest overall evaluation result in each of the two lots.

1.3.12. Variant bids (Nebenangebote) Variant bids are permitted for each lot, also without a main bid. A variant bid may deviate from the specification only in the following respects: (a) the start date, which may be later than the permitted window but no later than 1 February 2027 (Lot 1) or 1 May 2027 (Lot 2); (b) the distribution of GPU capacity over time, provided the binding minimum FLOPs are delivered within the base period and at least 1,024 GPUs (Lot 1) or 2,048 GPUs (Lot 2) are available concurrently at one site in one InfiniBand fabric for the main training phases; (c) additional capacity at further EU/EEA sites. All other requirements marked F, the price ceiling, the binding minimum FLOPs, the payment and guarantee terms (1.3.7) and the end of all services including options no later than 30 September 2027 are minimum requirements for variant bids. Variant bids must be clearly marked as such, with a separate requirements sheet and price sheet, and a description of each deviation. They are evaluated with the same award criteria as main bids (Section 2); criterion A4 applies one linear scale to main and variant bids.

1.4. Requirements

Page 8 of 28

[Seite 9]

The following criteria are applied to determine the most technically and economically advantageous offer for each lot:

• F — exclusion criteria. Failure to meet any requirement marked F results in the immediate rejection of the offer for the lot concerned (knock-out criterion).

• B — evaluation criteria, assessed on a points-based scale with the weighting given in Section 2.

• C — compliance criteria, to be answered or confirmed but not scored.

• I — informative criteria, optional additional information that does not influence the evaluation result.

All fields requiring bidder information must be completed in full for each lot bid for. Missing or insufficiently described functions, features, performance data or measurement methods in relation to exclusion criteria (F) result in exclusion from the lot. The restriction to NVIDIA reference architectures of the Blackwell generation or newer (DGX/HGX nodes for Lot 1, NVL72 racks for Lot 2) or technically equivalent systems (A1), and to InfiniBand interconnects, is a technical necessity derived from the workload and the validated software stack; the justification is given in Annex A. Where requirements differ between lots, both values are stated.

AGPU resources
No.ItemBidder informationNotesCrit.
A1GPU generation and system class( ) yes / ( ) no Bidder must specify the NVIDIA GPU model and system offered and its vendor-specified non- sparse peak BF16 throughput per GPU.Lot 1: GPU nodes according to the NVIDIA DGX/HGX B200 or B300 reference architecture (8 GPU modules per node, NVLink/NVSwitch) or newer NVIDIA reference architectures of the same class (e.g. DGX Rubin NVL8), or technically equivalent nodes meeting A7, A10 and A11. Lot 2: Rack-scale systems according to the NVIDIA GB200 NVL72 or GB300 NVL72 reference architecture or newer NVIDIA rack-scale reference architectures with a coherent NVLink domain of at least 72 GPUs (e.g. Vera Rubin NVL144), or technically equivalent rack-scale systems meeting A7, A10 and A11. Offers with NVIDIA Hopper (H100, H200), Ampere (A100) or older GPUs and GPUs from other vendors are excluded in both lots (Annex A.2 to A.4).FI
A2Dedicated cluster size( ) yes / ( ) no Bidder must specify the number of nodes or racks, the GPU count, the node or rack configurationLot 1: a dedicated cluster of at least 128 nodes (1,024 GPUs) during the first 65 days of the term and at least 16 nodes (128 GPUs) thereafter, available concurrently and exclusively to the contracting authority, without queueing,FI

Page 9 of 28

[Seite 10]

AGPU resources
No.ItemBidder informationNotesCrit.
and the NVLink domain size.oversubscription or contention with other customers; the contracting authority may call up to 128 nodes thereafter within the guaranteed FLOPs. The 128 nodes of the minimum profile must be located at one site in one InfiniBand fabric; additional capacity may be located at other EU/EEA sites. Lot 2: a dedicated cluster of at least 29 NVL72 racks (2,048 GPUs) from the start date until the guaranteed FLOPs are delivered. For GPUs with a higher peak BF16 throughput the count may be reduced proportionally (minimum count = reference count × reference peak / offered peak), rounded up to full nodes or racks.
A3Minimum guaranteed FLOPs( ) yes / ( ) no Bidder must declare the minimum guaranteed FLOPs delivered within the base period of the lot.Binding minimum 1.5 × 10²⁵ FLOPs for Lot 1 (≈1.85 million B200 GPU-hours) and 2.4 × 10²⁵ FLOPs for Lot 2 (≈2.67 million GB200 GPU-hours), computed as GPU-hours × vendor-specified non- sparse peak BF16 throughput per GPU. Offers below these values are excluded. For other admissible GPU models (A1) the equivalent is calculated from the vendor-specified peak BF16 throughput; bidders provide evidence that the cluster delivers the minimum within the period of performance (Section 1.2).F
A4Availability in the period of performance( ) yes / ( ) no Bidder must state the site(s), the number of racks or nodes and the start date.Is the complete cluster available from the start of the period of performance (1.3.2) and for its full duration, including the option months if exercised? The bidder declares bindingly the data centre site(s), the number of NVL72 racks (Lot 2) or GPU nodes (Lot 1) and the date from which the complete cluster is available (1.3.2). Evaluation (B): the offered start date is scored linearly between: Lot 1: 10 December 2026 or earlier (100%) to 1 February 2027 (0%); Lot 2: 15 January 2027 (100%) to 1 May 2027 (0%). TheFB

Page 10 of 28

[Seite 11]

AGPU resources
No.ItemBidder informationNotesCrit.
same scale applies to main bids and variant bids.
A5GPU service- level agreement( ) yes / ( ) noCluster-wide availability of at least 99.5% per month for compute nodes, NVLink, InfiniBand fabric and the shared storage used by training jobs. Jobs interrupted by provider-side failures are credited in GPU-hours; restart from the last checkpoint must be supported. Incident acknowledgement within 30 minutes, resolution or failover within 4 hours; hot spares: Lot 1 at least 3 spare nodes during the 128-node phase and at least 1 spare node thereafter; Lot 2 at least one spare NVL72 rack or equivalent compute trays. While not required for failover, the spare capacity may be used by the contracting authority for evaluation and debugging jobs. Planned maintenance announced 72 hours ahead, at most 4 hours per month in agreed windows. We define downtime as the period during which the reserved compute capacity (GPU cluster, CPU resources or shared storage) is unavailable to the contracting authority due to provider- side failures. The clock starts at detection or ticket opening and stops when the affected capacity is fully restored and usable. Planned maintenance meeting the conditions specified above does not count as downtime. Measurement methodology: Availability is measured per calendar month as 𝐴𝑣𝑎𝑖𝑙𝑎𝑏𝑖𝑙𝑖𝑡𝑦 = 1− 𝑇𝑜𝑡𝑎𝑙 𝑑𝑜𝑤𝑛𝑡𝑖𝑚𝑒 𝑚𝑖𝑛𝑢𝑡𝑒𝑠 ×100 𝑇𝑜𝑡𝑎𝑙 𝑚𝑖𝑛𝑢𝑡𝑒𝑠 𝑖𝑛 𝑚𝑜𝑛𝑡ℎ Penalties and crediting: • Any training job interrupted due to provider-side failures must be fully recredited in compute hours. • If monthly availability falls below the stated SLA thresholds (e.g., 99.5% for GPUs), the provider must recredit all lost compute hours plus an additionalFB

Page 11 of 28

[Seite 12]

AGPU resources
No.ItemBidder informationNotesCrit.
buffer of 10% to compensate for rescheduling overhead. • Persistent underperformance (≥ 3 consecutive months below SLA) may trigger contractual remedies up to termination. Evaluation (B): the bid with the highest SLA level receives 100%, other offers a percentage relative to it.
A6Local NVMe storage( ) yes / ( ) no Bidder must specify the NVMe configuration.Lot 1: at least 8 TB NVMe per node (F) for checkpoint staging (about 2 TB per full checkpoint), inference caches and temporary storage. Lot 2: at least 4 TB NVMe per compute tray (F) or an equivalent rack-level burst buffer able to stage a full checkpoint of up to 8 TB within about 3 minutes. Equivalent alternatives must be demonstrated without loss of performance or fault tolerance. Evaluation (B): Lot 1 ≥24 TB per node preferred, Lot 2 ≥8 TB per tray preferred; the best offer receives 100%, others proportionally.FB
A7NVLink domain( ) yes / ( ) noLot 1: all GPUs of a node are interconnected via NVLink/NVSwitch (NVLink 5 or newer, at least 1.8 TB/s per GPU) in one NVLink domain of at least 8 GPUs. Lot 2: coherent NVLink/NVSwitch domain of at least 72 GPUs per rack (NVL72), at least 1.8 TB/s per GPU. PCIe- only interconnect between GPUs is not acceptable. Bidders describe the NVLink domain size and topology.FI
A8Inter-node and inter-rack connectivity( ) yes / ( ) no Bidder must describe topology type, number of rails, blocking ratio, per- link and aggregateInfiniBand NDR (400 Gb/s per GPU), XDR (800 Gb/s) or faster, fully available across all GPU nodes or racks of the offered cluster. Ethernet-based fabrics (including RoCE, Spectrum-X, EFA or similar) do not qualify. Non-blocking (1:1), rail-optimised topologies areFI

Page 12 of 28

[Seite 13]

AGPU resources
No.ItemBidder informationNotesCrit.
bandwidth, measured latency for ≤8 KB and ≤1 MB messages, and SHARP support.strongly preferred; blocking ratios must be disclosed. Lot 2: GPU-initiated InfiniBand communication (IBGDA via NVSHMEM) must be available for expert- parallel all-to-all across racks. Justification: Annex A.1.
A9Collective communication( ) yes / ( ) no Bidder must confirm NCCL compatibility or describe a demonstrably equivalent library.Full compatibility with the NVIDIA Collective Communication Library (NCCL), or an equivalent library delivering the same functionality and performance for all-reduce, all-gather, all-to-all, reduce- scatter and broadcast across GPUs, nodes and racks. Lack of NCCL compatibility or equivalent functionality is an exclusion criterion.FI
A10GPU memory( ) yes / ( ) no Bidder must specify HBM capacity and bandwidth per GPU.At least 180 GB HBM per GPU with at least 7.5 TB/s memory bandwidth; Lot 1 at least 1.4 TB of GPU memory per 8-GPU NVLink domain, Lot 2 at least 13 TB of GPU memory per NVL72 rack. Derivation: Annex A.3 and A.4.FI
A11FP8 and FP4 Tensor Cores( ) yes / ( ) noThe GPUs must support FP8 and FP4 (NVFP4) Tensor Core arithmetic for training and inference, as implemented in the NVIDIA Blackwell architecture. GPUs without FP4 Tensor Cores are excluded. Derivation: Annex A.3.F
A12Node CPUs and memory( ) yes / ( ) no Bidder must specify core count and system RAM per GPU node or compute tray.Lot 1: at least 64 physical cores and 1 TB system RAM per GPU node (F); ≥112 cores and ≥2 TB preferred. Lot 2: NVIDIA Grace or successor (e.g. Vera) CPUs with at least 480 GB LPDDR5X per superchip, coherently connected to the GPUs (NVLink-C2C). Node CPUs do not count towards the dedicated CPU resources in Section B.FI
A13Software and job orchestration( ) yes / ( ) noFully managed (installation, upgrades and support included): • NVIDIA drivers, CUDA and associated CUDA libraries, including NCCL and NVSHMEM. • GPU/InfiniBand communication: Support for NVIDIA GPUDirectF

Page 13 of 28

[Seite 14]

AGPU resources
No.ItemBidder informationNotesCrit.
technologies, including GPUDirect RDMA and GPUDirect Async (IBGDA). • Cluster and workload management: SLURM. Kubernetes may be provided additionally where useful for the operation of containerised workloads. Jobs must be able to launch Ray clusters, inference servers and sandboxed environments inside SLURM allocations (asynchronous reinforcement learning with NeMo RL and NeMo Gym). • Container runtime: A GPU- and InfiniBand-capable container runtime supporting GPU and InfiniBand pass- through, such as Enroot/Pyxis, Apptainer, Singularity, Docker, or technically equivalent solutions. • NVIDIA NGC and AI software: Current compatible NVIDIA NGC container images and software environments, including, as applicable, PyTorch, NeMo, Megatron-Bridge, vLLM and TensorRT-LLM. • Programming and developer environment: Python 3.14 or newer, Git, GCC and pip (Python Package Index / PyPI access). • Monitoring and observability: Monitoring capabilities for system and workload operation, including, as applicable, GPU utilisation, CPU and memory utilisation, storage, network and InfiniBand fabric status, and NCCL communication/performance.
A14Performance benchmarksSheet with MLPerf Training resultsMLPerf Training results (version 5.0 or newer) for LLM-relevant workloads, in particular Llama 3.1 405B pretraining, DeepSeek-V3 (MoE) pretraining and Llama 2 70B LoRA, for the proposed GPU architecture and system class. Official MLPerf results for the proposed architecture must be cited where available; otherwise self-run MLPerf- compliant benchmarks with logs andFB

Page 14 of 28

[Seite 15]

AGPU resources
No.ItemBidder informationNotesCrit.
documentation, auditable by the contracting authority. Evaluation (B): the bid with the highest LLM-relevant MLPerf performance per GPU receives 100%, others proportionally.
BCPU resources, storage and data transfer
No.ItemBidder informationNotesCrit.
B1Dedicated CPU resources( ) yes / ( ) no Bidder must describe node configuration, scheduling in coordination with GPU jobs and container isolation.Lot 1: at least 15 million CPU core-hours over the base period as a dedicated partition of at least 4,096 physical x86-64 cores available concurrently. Lot 2: a dedicated partition of at least 8,192 physical cores (x86-64 or Arm). Both: ≥8 GB RAM per core, ≥4 TB local NVMe per CPU node. Used for tokenisation, deduplication and filtering of about 1 PB of text, orchestration, and the sandboxed execution environments of reinforcement learning and agentic data generation (Lot 2: ≥8,192 concurrently running isolated containers). CPU jobs run in the same SLURM environment and share the parallel filesystem with the GPU cluster. Minor variations of the project plan are possible, provided that the overall training workload is completed by the respective planned end date.FI
B2Shared storage( ) yes / ( ) no Bidder must specify filesystem, usable capacity and measured throughput.A POSIX parallel filesystem (e.g. Lustre, BeeGFS, Spectrum Scale, WEKA or equivalent) with at least 1 PB usable capacity (F), accessible concurrently from all GPU and CPU nodes, plus at least 1 PB of S3-compatible object storage (F) (or an additional 1 PB of parallel filesystem) for raw corpora, archives and data delivery. Preferred sustained throughput ≥150 GB/s write and ≥400 GB/s read; minimum floor and alternative patterns in B5.FB

Page 15 of 28

[Seite 16]

BCPU resources, storage and data transfer
No.ItemBidder informationNotesCrit.
Evaluation (B): offers with at least twice the minimum usable parallel filesystem capacity receive 100%, others proportionally (usable capacity ÷ twice the minimum).
B3CPU network connectivity( ) yes / ( ) noEach CPU node is connected to the shared storage and to the GPU cluster with at least 100 Gb/s.F
B4CPU service- level agreement( ) yes / ( ) noAvailability of at least 99.0% per month for the CPU partition; queueing delays of at most 24 hours under the reserved capacity; automatic restart of jobs within 1 hour after a node failure.F
B5Storage service-level agreement( ) yes / ( ) noAggregate sustained throughput of at least 50 GB/s write and 150 GB/s read (F) with about 2 GB/s per active compute node, measured with IOR/FIO under declared conditions (≥64 readers or ≥32 writers). Lot 2: a full checkpoint of up to 8 TB must be written within about 3 minutes (≥45 GB/s sustained from the GPU cluster). If the preferred ≥150/400 GB/s cannot be met, the bidder must demonstrate an alternative (node-local staging, dataset sharding, checkpoint compression, prefetching) with equivalent end-to-end training performance: no data-loading stalls and checkpoint/restart within about 3 minutes, evidenced by step-time measurements and a fail-and-restart demonstration. Storage availability ≥99.5% per month. Evaluation (B): the bid with the highest measured throughput receives 100%, others proportionally.FB
B6Data ingress and egress( ) yes / ( ) noExternal connectivity of at least 100 Gb/s. Lot 1: ingestion of about 1 PB (network transfer or physical media import), possible at the latest from the start of the test phase (G2). Lot 2: ingestion of about 150 TB of tokenised training data before the start date and continuous exchange ofF

Page 16 of 28

[Seite 17]

BCPU resources, storage and data transfer
No.ItemBidder informationNotesCrit.
checkpoints (up to about 8 TB each) with Lot 1. Ingress and egress free of charge (see C4).
CPricing model and billing options
No.ItemBidder informationNotesCrit.
C1Invoicing period( ) yes / ( ) noCan the full price of the base period be invoiced no later than 1 December 2026 and does the bidder accept payment in advance against the bank guarantee (1.3.7)?F
C2Bank guarantee for advance payment( ) yes / ( ) noIs a free bank guarantee for the advance payment possible, covering services not yet delivered at the time of payment (1.3.7)?F
C3Storage costs included( ) yes / ( ) noAre the storage costs (B2) included in the price?F
C4Data transfer included( ) yes / ( ) noAre ingress and egress of data (Lot 1: about 1 PB in, several hundred TB out; Lot 2: about 150 TB in, several hundred TB of checkpoints out) included in the price?F
C5Early data ingress( ) yes / ( ) noCan data be ingested into the storage (B2) already within two weeks after conclusion of the contract, i.e. before the test phase (G2) and the start date (1.3.2)?C
C6Options( ) yes / ( ) noDoes the bidder confirm the option of up to three months (1.3.2) under identical conditions (Lot 1: 128 GPUs, Lot 2: 2,048 GPUs) at the offered unit price per GPU- hour, ending no later than 30 September 2027? The unit price for option months is calculated by the contracting authority as the total price of the base period (all services included) divided by the guaranteed GPU-hours of item 1a of the price sheet; it covers the continued provision of all services during the option months (1.3.2).F

Page 17 of 28

[Seite 18]

DSustainability and data center location
No.ItemBidder informationNotesCrit.
D1Data center location( ) yes / ( ) no Bidder must specify the primary data center location(s) per lot.All compute services must be provided exclusively within the EU/EEA or in a third country with an adequacy decision by the European Commission under GDPR (Annex C).F
D2Power usage effectivenessPUECompliance and comparability only; not scored.C
D3Renewable energyDocumentationBidders state whether renewable energy is used; not scored.C
D4Sustainability reportsReportsBidders state whether sustainability reports are provided; not scored.C
ELegal, security, certifications and compliance
No.ItemBidder informationNotesCrit.
E1GDPR compliance( ) yes / ( ) noAre your services compliant with the GDPR?F
E2Data jurisdiction( ) yes / ( ) noSelect "yes" if the processed data is not subject to the US CLOUD Act, Patriot Act or similar third-country laws.F
E3EU data residency( ) yes / ( ) noDo you support EU data residency and processing requirements, including a contractual commitment to the named facility or facilities?F
E4CertificationsList in an extra sectionWhich certifications does the data center have (e.g. ISO/IEC 27001, ISO 50001)?I
E5Access control( ) yes / ( ) noDedicated tenancy, role-based access, multi-factor authentication and audit logs available?F
E6References: public institutionsUp to 3 referencesPrevious assignments under EU or national procurement regulations, with use cases and resource usage (maximum number of GPUs, job duration).I
E7References: LLM trainingUp to 3 referencesExperience with LLM training at a scale of 1,000 GPUs or more outside of MLPerf benchmarking; for Lot 2, experience with NVL72 rack-scale training.I

Page 18 of 28

[Seite 19]

FPrice (per lot)
No.ItemBidder informationNotesCrit.
F1aTotal guaranteed FLOPsBidder declares the minimum guaranteed FLOPs delivered within the base period of the lot.Binding minimum 1.5 × 10²⁵ FLOPs (Lot 1) / 2.4 × 10²⁵ FLOPs (Lot 2); offers below are excluded. Within a lot, the bid with the highest guaranteed FLOPs receives 100 points; other bids: Score = (Bidder’s FLOPs ÷ Highest FLOPs bid) × 100. For GPUs other than the reference model, the FLOP equivalent is computed from the vendor-specified peak BF16 throughput.FB
F1bEffective unit price in EUR per 10²³ FLOPs (incl. VAT)Bidder declares the total contract price for the base period of the lot in EUR (incl. VAT). The contracting authority calculates EUR per 10²³ FLOPs from price and guaranteed FLOPs.Maximum total contract value for the base period: 16,400,000 € (incl. VAT) (Lot 1), 26,100,000 € (incl. VAT) (Lot 2); offers above are excluded. Within a lot, the bid with the lowest EUR per 10²³ FLOPs receives 100 points; other bids proportionally relative to the lowest value.FB
GSupport, Setup and User Management (GPU and CPU)
No.ItemBidder informationNotesCrit.
G1Support and Cluster Operationsyes ( ) / no ( ) The bidder must describe the available support model, including the relevant expertise, availability and scope of services.Will you provide an expert cluster operations team that ensures the well- planned usage of the provided resources? Both lots: The bidder shall provide access to an expert cluster operations team for the planning, onboarding, operation, optimisation of the provided resources, error root cause analysis and debugging support. On-demand and regular support via remote sessions between the bidder and the Contracting Authority shall be available throughout the contract period. Upon request by the Contracting Authority, the support shall include, as applicable, planning sessions covering data ingress and egress, network configuration, CPU and GPU utilisation, storage capacity and utilisation, as well asF

Page 19 of 28

[Seite 20]

GSupport, Setup and User Management (GPU and CPU)
No.ItemBidder informationNotesCrit.
dedicated onboarding and offboarding sessions and regular service review sessions. This requirement is in addition to the SLA requirements.
G2Test Phaseyes ( ) / no ( )Will you offer the possibility of a test phase in the target environment on a subset of the resources before start of the full resource usage? The bidder shall provide a test phase in the target environment using a subset of the contracted resources prior to the start of full resource utilisation. Lot 1: The test phase shall be available at least two weeks prior to the start of full resource utilisation. Lot 2: The test phase shall be available at least two weeks prior to the start of full resource utilisation. The test phase shall be conducted using the same relevant hardware, software, network and access environment as the target environment, to the extent applicable.F
G3User Accountsyes ( ) / no ( ) The bidder must describe any applicable technical or organisational limitations concerning the number of users or user accounts.Will you provide access to the provided resources for up to 100 users from the contracting authority, its project partners and associated partners? Both lots: The bidder shall provide access to the contracted resources for up to 100 individual users from the contracting authority, its project partners and associated partners. The solution shall support the creation, management and removal of individual user accounts throughout the contract period (e.g., via LDAP).F
G4Login Node Capacity( ) yes / ( ) no The bidder shall specify the CPU model, number of physical CPU cores / threads and availableBoth lots: The bidder shall provide login nodes with sufficient CPU and memory resources to support the concurrent access and interactive workloads of the expected user base without the login nodes becoming a bottleneck for normal cluster operations. The login nodes shallF

Page 20 of 28

[Seite 21]

GSupport, Setup and User Management (GPU and CPU)
No.ItemBidder informationNotesCrit.
RAM per login node, as well as the number of login nodes provided.provide sufficient CPU thread capacity and RAM for typical activities such as job preparation and submission, compilation, interactive sessions, software environment management and other non-compute-intensive user activities.

Page 21 of 28

[Seite 22]

2. Evaluation Criteria

Each lot is evaluated separately. The contracting authority has set fixed budget ceilings of 16,400,000 € (incl. VAT) (Lot 1) and 26,100,000 € (incl. VAT) (Lot 2) for the base periods. Bids exceeding the ceiling of a lot or falling below its binding minimum FLOPs are excluded before scoring. Since the purpose of this procurement is to maximise delivered compute within the available budget, the price evaluation is split into two sub-criteria that together account for 60% of the overall score: F1a (40%) total guaranteed FLOPs, and F1b (20%) effective unit price per 10²³ FLOPs. Functionality criteria (40%) are assessed separately and do not overlap with the FLOP evaluation.

No.Criterion blockWeight
1Price F1 (60%) F1a — Total guaranteed FLOPs (40%) F1b — Effective unit price per 10²³ FLOPs (20%)Price evaluation
2Functionality scope (40%) Bids within a lot are evaluated relative to each other: the best offer for a criterion receives 100 points, others proportionally on a linear basis. A4 Start date (absolute scale, see 1.4) — 15% A5 GPU service-level agreement — 5% A6 Local NVMe storage — 5% A14 Performance benchmarks (LLM-relevant MLPerf Training) — 5% B2 Shared storage capacity — 5% B5 Storage service-level agreement — 5%Functionality evaluation

Overall score per lot = price subtotal (max 60) + functionality subtotal (max 40). The offer with the highest overall score in a lot is the most economically advantageous tender for that lot. A worked example of the price evaluation is given in Annex B.


Place, Date Signature / Company Stamp

Page 22 of 28

[Seite 23]

3. Annex

3.1. Annex A — Technical Justification of the Hardware Requirements

3.1.1. InfiniBand Interconnect Expert-parallel mixture-of-experts training performs an all-to-all exchange of activations in every layer; its latency and tail behaviour directly determine step time. Asynchronous reinforcement learning transfers updated policy weights (several hundred GB in BF16 for models larger than 100 billion parameters) from training to inference workers several times per hour while rollouts are in flight, and streams rollouts back; jitter in this path stalls thousands of GPUs. Both patterns require a lossless, low-latency fabric with deterministic scaling across the whole cluster. InfiniBand provides in-network reductions (SHARP), credit-based lossless flow control, multi-rail non-blocking topologies and a dedicated training fabric isolated from storage and management traffic, and it is the interconnect for which the frameworks used by the contracting authority (NCCL, Megatron-Core, NeMo RL, vLLM) are validated at this scale. Ethernet-based fabrics (RoCE, Spectrum-X, EFA) rely on congestion-control tuning that is not validated for these workloads at 2,000+ GPUs, and their qualification is out of scope for a time-bound public procurement. The InfiniBand requirement is therefore a project-critical technical condition, not a vendor preference.

In the NVL72 rack-scale systems of Lot 2, the 72 GPUs of a rack communicate through the NVLink switch fabric of the rack; between racks, the GPUs are connected through their ConnectX-8 network interfaces at 800 Gb/s to an InfiniBand fabric (Quantum-X800), which is also the interconnect of NVIDIA’s own reference architecture for GB200 and GB300 NVL72 systems (DGX SuperPOD). The InfiniBand requirement in A8 therefore refers to the inter-node fabric of Lot 1 and to the inter-rack fabric of Lot 2; within a rack, NVLink applies (A7).

3.1.2. NVIDIA Software Stack The training pipeline of the contracting authority (CUDA, NCCL, PyTorch, Megatron-Core/Megatron- Bridge, Transformer Engine with FP8 and NVFP4 recipes, NeMo RL with vLLM) is validated and supported only on NVIDIA hardware. The models are mixture-of-experts architectures with state- space-model components whose reference implementations and training recipes (NVIDIA Nemotron 3 family [1], [2], [3]) are published for this stack. Porting and validating the full stack (state-space- model and MoE kernels, optimizer kernels, communication libraries, RL infrastructure) to another accelerator platform is infeasible within the project timeline. The restriction to NVIDIA GPUs is a technical necessity to ensure feasibility, reliability and timely delivery.

3.1.3. Blackwell Generation for Lot 1 and Exclusion of Hopper, Ampere and Older GPUs Lot 1 requires GPU nodes of the NVIDIA DGX/HGX reference architecture of the Blackwell generation or newer, with the DGX/HGX B200 platform as the minimum baseline (8 GPUs with 180 GB HBM3e each, 1.44 TB GPU memory per node, NVLink 5 with 1.8 TB/s per GPU, 2.25 × 10¹⁵ FLOPs/s dense BF16, 4.5 × 10¹⁵ FP8 and 9 × 10¹⁵ FP4 per GPU [11]). Four workloads of Lot 1 cannot be executed efficiently, or at all, on the preceding Hopper generation:

  1. Synthetic data generation with a 400B-class MoE generator. The generator holds about 400 GB of FP8 weights. On an 8 × B200 node (1.44 TB) about 1 TB remains for the key-value caches of long documents (context of 65K tokens, more than 1,000 concurrent sequences),

Page 23 of 28

[Seite 24]

the configuration on which the contracting authority’s throughput figures were measured [4], [13]. On 8 × H100 (640 GB) only 240 GB remain, forcing 16-GPU replicas across two nodes; independent benchmarks show 5 to 7 times higher FP8 generation throughput per GPU on B200 than on H100 for MoE models of this class (5,215 versus 739 tokens/s per GPU for a 671B MoE model) [6], so the same budget buys two to three times more data on Blackwell despite the higher hourly price.

  1. Ablation runs and evaluations in NVFP4 numerics. The model of Lot 2 is pretrained in NVFP4 4-bit floating point, as are NVIDIA Nemotron 3 Super and Ultra [1], [2], [10]. Data- mixture and numerics ablations for this recipe and the evaluation of NVFP4-quantised checkpoints are only meaningful with the same numerics. FP4 Tensor Cores exist on Blackwell GPUs only; Hopper supports FP8 at most.

  2. Reinforcement learning with verifiable rewards (RLVR) on models larger than 100B parameters. Each step generates 8,192 rollouts of up to 64K tokens with inference workers while the policy (several hundred GB in BF16), the reference model and the caches stay resident [1], [2]. Measured on B200, the reinforcement learning stage occupies 1,024 GPUs for about 30 days [4]; with half the memory and 40% of the bandwidth per GPU, Hopper would not complete this stage within the period of performance.

  3. One validated software stack. All throughput figures underlying this tender (pretraining, post-training and data generation) were measured on B200 [4], [13]; NVFP4 pretraining, asynchronous GRPO and vLLM with multi-token-prediction speculative decoding are published for and validated on Blackwell systems [1], [2], [5]. Operating and validating a second, older fleet in parallel is not possible within the project timeframe.

The comparison below (dense values from the vendor datasheets [11]) summarises the gap. The Blackwell requirement is an objective, workload-derived necessity and not a preference for the newest hardware; it is formulated by capability (memory, numerics, interconnect), so that newer generations qualify automatically.

Specification (dense,A100 SXMH100 SXMH200 SXMB200
vendor datasheet)(HGX/DGX)
HBM capacity80 GB80 GB141 GB180 GB
HBM bandwidth2.0 TB/s3.35 TB/s4.8 TB/s8 TB/s
BF16 Tensor Core (dense)312 TFLOPs/s989.5 TFLOPs/s989.5 TFLOPs/s2,250 TFLOPs/s
FP8 Tensor Core (dense)not supported1,979 TFLOPs/s1,979 TFLOPs/s4,500 TFLOPs/s
FP4 Tensor Core (dense)not supportednot supportednot supported9,000 TFLOPs/s
NVLink bandwidth per GPU600 GB/s900 GB/s900 GB/s1,800 GB/s
400B-class FP8 generator: memory left for caches per 8-GPU nodeFP8 not supported240 GB730 GB1,040 GB

3.1.4. NVL72 Rack-Scale Systems for Lot 2

Page 24 of 28

[Seite 25]

Lot 2 requires rack-scale systems of the NVIDIA GB200 or GB300 NVL72 reference architecture or newer with a coherent NVLink domain of at least 72 GPUs. The Lot 2 model (mixture-of-experts, more than 350 billion parameters) has up to about 8 TB of model and optimizer states per checkpoint and is trained with wide expert parallelism: in every MoE layer, tokens are exchanged all- to-all between the GPUs holding the experts. NVIDIA trained the reference model of this class (Nemotron 3 Ultra, 550B/55B) on GB200 NVL72 with the expert-parallel group mapped inside the NVLink domain [2]. According to our information, the 20-trillion-token run of this model takes about four to five weeks on 6,000 GB200 GPUs with current software. On 8-GPU nodes the same expert- parallel traffic crosses InfiniBand and dominates the step time [14]; NVIDIA’s published training benchmarks show 1.26 to 1.5 times higher throughput per GPU on GB200 than on B200 for large MoE models at equal GPU count [5]. The same holds for the rollout generation of reinforcement learning on the Lot 2 model, where independent inference benchmarks report 1.4 to 4 times higher throughput per GPU on GB200 NVL72 than on B200 at typical interactivity [15]. The market consultation preceding this tender (September 2026) confirmed availability of NVL72 capacity of the required scale and excluded Hopper systems for this workload.

The requirement is formulated by capability (coherent NVLink domain of at least 72 GPUs, at least 180 GB HBM per GPU, FP4 Tensor Cores), so that GB300 NVL72 and future rack-scale systems qualify. Bidders offering such systems compute the FLOP equivalent from the vendor-specified peak BF16 throughput (A3).

3.2. Annex B — Example of Price Evaluation (F1a + F1b), Lot 2

Assumptions: budget ceiling of 26,100,000 € (incl. VAT); binding minimum 2.4 × 10²⁵ FLOPs. F1a (40%): Score = (Bidder FLOPs ÷ Highest FLOPs) × 100. F1b (20%): €/10²³ FLOPs = (Total € ÷ FLOPs) × 10²³; Score = (Lowest €/10²³ ÷ Bidder €/10²³) × 100. Price subtotal (max 60) = 0.4 × F1a + 0.2 × F1b. The same method applies to Lot 1 with its ceiling and minimum.

BidderGuaranteed FLOPsTotal price (incl. VAT)€/10²³ FLOPsF1a (40%)F1b (20%)Price
subtotal
(max 60)
A2.8 × 10²⁵26,100,000 €93,214100.0098.3459.67
B2.6 × 10²⁵25,000,000 €96,15492.8695.3356.21
C2.4 × 10²⁵22,000,000 €91,66785.71100.0054.28

Calculations:

Bidder ABidder BBidder C
€/10²³ FLOPs(26,100,000 ÷ 2.8 × 10²⁵) × 10²³ = 93,214(25,000,000 ÷ 2.6 × 10²⁵) × 10²³ = 96,154(22,000,000 ÷ 2.4 × 10²⁵) × 10²³ = 91,667
F1a (reference 2.8 × 10²⁵)100.002.6/2.8 × 100 = 92.862.4/2.8 × 100 = 85.71
F1b (reference 91,667)91,667/93,214 × 100 = 98.3491,667/96,154 × 100 = 95.33100.00

Page 25 of 28

[Seite 26]

Price subtotal (F1) 0.4 × 100 + 0.2 × 98.34 = 0.4 × 92.86 + 0.2 × 0.4 × 85.71 + 0.2 × 59.67 95.33 = 56.21 100 = 54.28

Result: Bidder A wins the price block because the 40% weight on total FLOPs favours the largest compute delivered within the ceiling.

3.3. Annex C — Data Protection and Location Requirements

The requirement that all services be delivered from data centers located in the EU/EEA or in third countries with an adequacy decision by the European Commission arises directly from the General Data Protection Regulation (GDPR, Regulation (EU) 2016/679) and its interpretation by the Court of Justice of the European Union in the Schrems II judgment (C-311/18, 16 July 2020). GDPR Articles 44–49 regulate transfers of personal data to third countries; such transfers are only permissible if an adequacy decision exists (Article 45) or, in specific cases, if appropriate safeguards are in place (Articles 46–49). In Schrems II the CJEU invalidated the EU–US Privacy Shield and emphasised that contractual or technical measures may not be sufficient to compensate for systemic risks in certain jurisdictions. Public authorities in the EU may therefore not lawfully contract services that involve data processing in countries without an adequacy decision unless exceptional derogations apply, which is not the case for large-scale, continuous LLM training data; the European Data Protection Board has repeatedly confirmed this in its guidance. The location requirement in D1 is consequently a mandatory exclusion criterion (F).

3.4. Annex D — Supporting Evidence for Technical Requirements

All quantitative requirements of this tender are derived from the target workload (Sections 1.1 and 1.2) and from measurements on the same model family. The chain from compute to storage is summarised below; the references are listed in 3.4.2.

3.4.1. Compute-to-Storage Causal Chain

RequirementDerived valueRationale / supporting evidence
GPU memory≥180 GB per GPU; Lot 2 ≥13 TB per rackPolicy weights of several hundred GB in BF16, reference and teacher models, optimizer states and 64K-token rollout caches resident during RLVR; 400B-class FP8 generator with about 1 TB of caches per node; up to 8 TB of model states per Lot 2 checkpoint [1], [2], [4].
NVLink domainLot 1 ≥8 GPUs; Lot 2 ≥72 GPUsExpert-parallel all-to-all of the Lot 2 model inside the NVLink domain [2], [5], [14]; rollout generation for the Lot 2 model [15].
NumericsFP8 and FP4 (NVFP4) Tensor CoresNVFP4 pretraining recipe of the Lot 2 model and matching ablations [1], [2], [10]; FP8 inference for data generation and rollouts [6], [13].

Page 26 of 28

[Seite 27]

RequirementDerived valueRationale / supporting evidence
InterconnectInfiniBand NDR/XDR, non- blocking preferred; IBGDA for Lot 2Expert-parallel all-to-all across nodes and racks; weight synchronisation between training and inference workers in asynchronous RL [1], [5], [7], [9], [12].
Dataset size and transfer≈1 PB raw text (Lot 1); ≈150 TB tokenised data into Lot 2; ≥1 PB parallel filesystem plus ≥1 PB object storageWeb, synthetic and translated corpora; training tokens at 4 bytes per token; checkpoints of the Lot 1 model (about 2 TB each) and of the Lot 2 run (up to 8 TB each, several dozen retained).
Checkpoint sizeabout 2 TB (Lot 1) and up to 8 TB (Lot 2) per full checkpointParameters × 14 bytes (BF16 weights, FP32 master weights, two Adam moments) [7]; consistent with a checkpoint size of about 7 TiB for the reference model according to our information.
Checkpoint throughput≥150 GB/s sustained writes preferred; Lot 2 ≥45 GB/s floorAn 8 TB checkpoint completes in about 55 seconds at 150 GB/s and in about 3 minutes at 45 GB/s [7], [8].
Dataset streaming≥400 GB/s sustained reads (preferred)Continuous feeding of 2,048 GPUs without starvation, in line with published HPC/MLSys storage analyses [7], [8].
Local NVMeLot 1 ≥8 TB per node; Lot 2 ≥4 TB per tray or rack- level burst bufferStaging of 2 to 3 checkpoints, inference caches, temporary storage; DGX B200 reference 8 × 3.84 TB [11].
CPULot 1 ≥15 million core- hours, ≥4,096 cores; Lot 2 ≥8,192 coresTokenisation, deduplication and filtering of ~1 PB; ≥8,192 concurrent sandboxes for RL environments (8,192 rollouts per RL step) [1], [2], [12].

3.4.2. Reference List [1] NVIDIA (2026). Nemotron 3 Super: Open, Efficient Mixture-of-Experts Hybrid Mamba- Transformer Model for Agentic Reasoning. arXiv:2604.12374. https://arxiv.org/abs/2604.12374

[2] NVIDIA (2026). Nemotron 3 Ultra: Open, Efficient Mixture-of-Experts Hybrid Mamba- Transformer Model for Agentic Reasoning. arXiv:2606.15007. https://arxiv.org/abs/2606.15007

[3] NVIDIA (2025). Nemotron 3 Nano: Open, Efficient Mixture-of-Experts Hybrid Mamba- Transformer Model for Agentic Reasoning. arXiv:2512.20848. https://arxiv.org/abs/2512.20848

[4] Contracting authority (2026). Internal throughput measurements on NVIDIA B200 systems for pretraining, post-training and synthetic data generation of the model family concerned (September 2026).

Page 27 of 28

[Seite 28]

[5] NVIDIA (2026). Megatron-Bridge Performance Summary, NeMo containers 25.09 to 26.04: DGX B200, GB200 NVL72 and GB300 NVL72 training throughput for DeepSeek-V3, Qwen3- 235B, Nemotron 3 Super and Llama 3.1 405B. https://docs.nvidia.com/nemo/megatron- bridge/latest/performance-summary-archive.html

[6] SemiAnalysis InferenceX (2026). DeepSeek-R1 inference throughput, B200 vs. H100 (FP8, 8K input / 1K output, 45 tokens/s per user: 5,215 vs. 739 tokens/s per chip). https://inferencex.semianalysis.com/compare/deepseek-r1-b200-vs-h100

[7] Narayanan, D., et al. (2021). Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM. SC 2021. https://arxiv.org/abs/2104.04473

[8] MLCommons (2025–2026). MLPerf Training results, v5.0 to v6.0 (Llama 3.1 405B, DeepSeek- V3, Llama 2 70B LoRA and other workloads). https://mlcommons.org/benchmarks/training/

[9] NVIDIA (2024). Nemotron-4 340B Technical Report (training on 768 DGX H100 nodes, 6,144 GPUs, InfiniBand). arXiv:2406.11704. https://arxiv.org/abs/2406.11704

[10] NVIDIA (2025). Pretraining Large Language Models with NVFP4. arXiv:2509.25149. https://arxiv.org/abs/2509.25149

[11] NVIDIA. DGX B200, GB200 NVL72 and GB300 NVL72 datasheets; H100, H200 and A100 datasheets. https://www.nvidia.com/en-us/data-center/dgx-b200/ ; https://www.nvidia.com/en-us/data-center/gb200-nvl72/ ; https://www.nvidia.com/en- us/data-center/h100/ ; https://www.nvidia.com/en-us/data-center/a100/

[12] NVIDIA (2026). NeMo RL and NeMo Gym documentation (asynchronous GRPO, vLLM generation workers, sandboxed environments). https://docs.nvidia.com/nemo/rl/latest/ ; https://docs.nvidia.com/nemo/gym/

[13] KletterMix (2026). German pretraining corpus of about 725 billion tokens translated with Qwen3.5-397B-A17B-FP8 on 126 nodes × 8 NVIDIA B200 in about ten days (Appendix A.2). arXiv:2606.03773. https://arxiv.org/abs/2606.03773

[14] PyTorch (2026). Enabling up to 41% faster pre-training with MXFP8 and DeepEP for DeepSeek-V3 on B200 with TorchTitan (inter-node all-to-all dominates step time at EP=32). https://pytorch.org/blog/enabling-up-to-41-faster-pre-training-mxfp8-and-deepep-for- deepseek-v3-on-b200-with-torchtitan/

[15] SemiAnalysis InferenceX (2026). GB200 NVL72 vs. B200: disaggregated DeepSeek-R1 FP4 inference (1.35 to 4.4 times higher throughput per GPU at 75 to 125 tokens/s per user). https://inferencex.semianalysis.com/blog/gb200-nvl72-vs-b200-disagg-deepseek-r1-fp4- dynamo-trt

Page 28 of 28

Alle Unterlagen dieser Ausschreibung