All notes

Archive of aionprem.pl. Newest first.

32 notes
Choosing a GPU and quantization when the H200 is unavailable
[on-prem]

The H200 rarely arrives on demand, and the rollout decision does not wait for delivery. We break down GPU selection for a 70B model from the availability side: how much memory it eats per precision, what it fits on after FP8 and INT4 quantization, and which card to take instead of the H200.

What a million tokens costs on-prem: H100 vs H200 vs API
[architecture]

Peak throughput from a benchmark does not tell you what a million tokens costs. The real number is the card's hourly rate divided by tokens per hour, then by utilization. We price a million tokens on H100 and H200, compare with API and show the utilization above which your own GPU wins.

AI TCO calculator: methodology, a worked example and FAQ
[vendor-evaluation]

An AI TCO calculator answers one question: at what utilisation does on-prem beat cloud. Methodology, three formulas (per token, GPU rental, on-prem CAPEX), a full 3-year worked example and FAQ. Verdict: at moderate volume the API wins on price, on-prem earns its keep on control, not cost.

AI Vendor DPA: 8 Clauses Whose Absence Breaks Your Audit
[compliance]

A vendor's boilerplate DPA stays silent exactly where an auditor looks first. Eight clauses whose absence breaks an NIS2 or GDPR audit: from the sub-processors behind the model API and training use of your data, to logs, breach notice and data deletion.