My answer to the reader asking whether AI spending is justified is this: for ordinary service workloads, direct the next increment first toward usable workflows and lower review costs. Shared compute and domestic models have important roles. A hardware count or a model ranking, however, is insufficient evidence for the next expansion.
Following the earlier explanation of budget scope, I tested that judgment with a deliberately narrow calculation. Instead of assigning one return on investment to a national budget, ask: how busy must continuously rented GPUs be to beat a metered API for the same model family and request length? Then ask what that saving buys in human review time.
The result has two layers. Dedicated rental breaks even at 33.5% useful busy time in a favorable baseline; a performance-and-overhead scenario raises the threshold to 59.8%. Yet even in a case where rental clearly saves money, the saving equals only 0.11 extra review seconds at an assumed wage. With equal acceptance rates, that small review-time difference could reverse the price advantage. Infrastructure efficiency matters. The calculation helps identify which efficiency matters most.
Separate published facts from assumptions
Public prices were checked on 6 October 2026. All amounts are US dollars, excluding tax.
- Prices: Lambda lists its two-H100 SXM plan at $4.19 per GPU-hour. Together lists Llama 3.3 70B Instruct Turbo FP8 at $1.04 per million input tokens and the same output price. Lambda · Together
- Performance: NVIDIA's NIM 1.8.0 FP8 TP2, two-H100 benchmark reports 3,342.67 system output tokens/second for 1,000 input and 1,000 output tokens at concurrency 150; TTFT is 428.54ms and ITL 44.44ms. The hardware was DGX H100, not a measured Lambda deployment. Results · Metrics
- Eligibility: Matching the model family and FP8 does not establish identical checkpoints, quality or service levels. Korean is absent from the model's explicitly supported languages. This is no recommendation for Korean public-service deployment. Meta model card
- My scenarios: Retaining 70% of published throughput, adding 25% to rental cost, and paying reviewers $30/hour are analytical assumptions, not observations.
This compares continuous cloud rental with metered inference, not national GPU purchases. Ownership requires a different total-cost calculation covering construction, power, cooling, staffing, replacement, residual value and contract horizon. Multiplying the percentages here by a national budget would exceed what the model establishes.
The 33.5% threshold asks for evidence of demand
The reference request costs $0.00208 through the API. Two rented GPUs cost $8.38/hour. If published performance transfers perfectly and additional operating cost is zero, dedicated cost is approximately $0.00069638 divided by u, the useful busy fraction of paid time.
The denominator of u is time paid for; its numerator is time spent doing useful work at the specified load. It is not the utilization number on a GPU dashboard. A machine producing unusable answers or repeated failed attempts can look busy. A thin stream of arrivals spread across a day also differs from a dense queue processed during a shorter window.
| Condition | Dedicated cost per 1,000 requests | Comparator |
|---|---|---|
| Baseline, 25% useful busy time | $2.786 | API $2.080 |
| Baseline, 60% useful busy time | $1.161 | API $2.080 |
| 70% throughput, 25% extra cost, 60% busy | $2.073 | API $2.080 |
The baseline breaks even at 33.5%. The performance-and-overhead scenario requires 59.8%. The last row matters: “60% is enough” cannot become a procurement rule. A small change in conditions can consume almost the entire saving.
How useful busy time changes serving cost
Vertical axis: USD per 1,000 attempted requests. Horizontal axis: useful paid-time busy fraction.
- A · solid teal line · Dedicated rental, ideal scenario: full reference throughput and no extra operating cost.
- B · solid orange line · Dedicated rental, assumed 70% throughput retention and 25% additional operating cost.
- C · dashed dark line · API baseline for the reference request.

Calculated comparison values, approximately
- C · API cost per 1,000 requests
- $2.08
- A · Ideal dedicated break-even
- 33.5% useful busy time
- B · Adjusted dedicated break-even
- 59.8% useful busy time
- Useful busy time is paid time producing useful work, not a GPU dashboard utilization reading.
- Conditional calculations; performance transfer, equal task quality and service equivalence have not been demonstrated.
Separating idle time from busy-period performance is an approximation. In a real service, declining demand can also lower concurrency and throughput. The deployment may move to a different curve altogether. The evidence procurement needs is an arrival record: which jobs come in, when, in what volume, and which can actually be pooled.
A quick first token does not mean a quick completed task
The reference timing implies roughly 44.8 seconds to receive the full output. Starting within a second and finishing a long answer within a second are very different promises. Published average-style metrics also cannot establish p95 or p99 performance under bursts.
Set the user's completion-time requirement first, then calculate cost using throughput that meets it. Test the API under the same requirement. This analysis did not measure which service is faster for a Korean workload.
Nor should short classification, document summarization and long-contract review be added together as interchangeable requests. Reading longer inputs and generating longer answers change a machine's capacity. Pooling different workloads may improve shared-compute economics; security partitions and response-time requirements may limit that pooling.
A fair batch comparison gives both sides permission to wait
Together specifies a 50% discount for this model's asynchronous Batch service, with a best-effort 24-hour target. That is a separate service condition. Batch documentation
I therefore also relax the dedicated-side latency screen, using the published concurrency-250 row at 4,167.39 output tokens/second. It is not a measurement of an optimally tuned dedicated batch system. Benchmark
| Service condition | Baseline break-even | 70% throughput and 25% extra cost |
|---|---|---|
| Reference comparison | 33.5% | 59.8% |
| Both sides allowed a relaxed-latency batch comparison | 53.7% | 95.9% |
Halving only the API price while freezing dedicated throughput would produce a more dramatic result, but it would not be a fair batch conclusion. The useful lesson is that deferrable work requires a joint comparison of price and scheduling. The ability to assemble a queue and keep capacity productively occupied becomes decisive.
When extra review reverses the ranking
At 60% useful busy time, baseline rental saves approximately $0.000919 per request. At a hypothetical $30/hour review wage, that buys 0.11 seconds. One review second costs $0.00833 under the same assumption. Across 1,000 requests, the infrastructure saving is about $0.92; one additional review second per request costs about $8.33.
When additional review costs outweigh the saving
Horizontal axis: USD per 1,000 attempted requests. Both bars share the same zero-based scale.
- A · teal bar · Ideal dedicated serving-cost saving versus the API at 60% useful paid-time busy fraction.
- B · blue bar · Cost of one additional review second per request for dedicated serving versus API, at a hypothetical $30/hour wage.

Conditional calculated values, approximately
- A · Serving saving per 1,000 requests
- $0.92
- B · Additional one-second review cost per 1,000 requests
- $8.33
- Additional review time that offsets the serving saving
- 0.11 seconds per request
- This is a comparison of a serving saving with an additional review cost, not two total service costs.
- Acceptance rates are assumed equal. Identical review costs cancel and leave the serving saving unchanged.
- The additional review time and wage are hypothetical; no actual review-time difference was measured.
This threshold is a difference in average review time between the approaches, assuming equal acceptance rates. If both require identical review, the common cost cancels and the dedicated system retains its absolute serving saving. No additional review burden was observed on either side. If acceptance rates also differ, compare their full accepted-task costs below.
This can strengthen the case for a domestic model. If it better fits a local task and reduces checking or correction, a higher token price may still deliver lower overall cost. Conversely, a cheap model's subtle errors can erase its serving advantage. The appropriate unit of evaluation is total cost per accepted task after review.
A simple accounting expression is:
Cost per accepted task = (inference + review seconds × hourly wage / 3,600 + other workflow costs) / acceptance rate.
This divides the total costs of evaluated attempts by accepted outcomes under one consistent rubric. It does not assume independent retries eventually succeed. Safety and legal eligibility remain prerequisites that an attractive average cost cannot offset.
Model development can face a related test. Suppose development consumes 25% of an otherwise equal short-run resource envelope, leaving the rest for serving. With equal acceptance rates, serving-and-review cost per attempted task must fall by more than 25% to produce more accepted work within that envelope. If the acceptance-rate ratio improves to 1.10, that per-attempt cost must be below 82.5% of the old level. This is an accounting hurdle, not a forecast that development achieves those improvements. A project justified by long-run research value needs that value evaluated separately.
The strongest counterargument is the capability an API does not sell
Using this serving calculation to reject research GPUs or domestic model development would compare the wrong products. Researchers may need to change weights, train new architectures, inspect internals, control protected-data workflows or restore service after an external supplier disappears. A completion API need not provide those rights.
Research infrastructure and trained operators can also generate knowledge that benefits people beyond the initial purchaser. Rejecting every project with low current inference demand could destroy an opportunity to learn capabilities needed later. A bounded investment in research and control is justified now when its purpose and users are credible. Perfect certainty is not a prerequisite for action.
But the alternatives must be broader than completion APIs. Compare ownership with rented raw compute, reserved capacity, research vouchers, domestic private supply, portable software and operator training. When equivalent rights and capability can be purchased more cheaply, ownership itself has a weaker claim to a premium.
Treating sovereignty as insurance also requires specifying the event and protection. Locally installed foreign hardware can remain dependent on components, software and power. Compare the annual premium and additional preventable loss against the same best feasible alternative over a matching period; publish recovery objectives and exercise results. National security cannot be reduced to one expected-dollar figure, but the word “sovereignty” should not dispense with a premium, coverage definition or renewal test.
Connect the investments and give each a different expansion test
My preferred sequence is:
- Start ordinary deployment with competitive procurement and small real-work trials. Record arrivals, blind Korean-language evaluations, review time and repeat use. If data preparation or workflow design is blocking adoption, address it first.
- Expand shared compute against demonstrated demand and access problems. Measure queue time, unused allocations, access for smaller teams and additional research that feasible private alternatives would not enable. A verified queue and lower service-equivalent total cost strengthen the expansion case.
- Evaluate domestic models for capabilities that are hard to substitute. Do they reduce review costs, enable otherwise unavailable work, or demonstrate modification, portability and recovery? A general benchmark ranking supplies only part of that answer.
These investments complement one another. Without usable tasks, compute can sit idle; without access to compute, prepared research can wait. An equal split or a confidently stated optimal percentage has no evidential basis until the binding constraint is measured.
The observations that would change my view are clear. Valuable queues plus lower total cost at the same quality, tail latency and resilience would move me toward more shared capacity. A domestic model that sufficiently reduces review cost or demonstrates irreplaceable control would strengthen the development case. Idle hardware and weak repeat usage would favor delaying the next purchase and changing contracts or operations.
The next won of AI spending deserves a stricter test. Give more funding to the proposal that establishes whether the next bottleneck is compute, usable work or trustworthy accuracy. On the evidence available here, measuring the workflow and review bottlenecks is the strongest starting point.
Reproduce the calculation
The public companion (ZIP) contains equations, executable code, selected-input CSV and source URLs. Replace performance, labor and review assumptions with measurements to find where the ranking reverses. This article estimates neither Korean procurement total cost nor actual Korean-language quality, national-program ROI or an optimal budget allocation.
Related reading: Distinguishing the scopes of AI budget figures. Establish what a budget number covers before applying a serving-cost threshold or a separate capability test.