AI Engineering Lab

Yogesh Kuchimanchi · Four small tools for inspecting models, retrieval, evaluation, and inference costs.

Use short, non-sensitive examples. Public calculators do not call a paid model. Results describe the selected input or recorded experiment.

Plan an inference workload

Choose a baseline, adjust inputs, then compare listed monthly costs.

Workload preset

Illustrative baselines; adjust to measured traffic and model needs.

Managed API Workload

Traffic and token inputs drive the API estimates below.

Self-Hosted Compute Setup

Allocated hours drive rental cost; model size drives memory fit.

Operational schedule
0 744
Precision (FP16 / INT8 / INT4)

VRAM estimate = (weights + KV-cache allowance) × 1.20 (decimal GB). The 20% buffer is a planning allowance, not measured activation memory. Set KV-cache for your context length and concurrency; these are not inferred from the API token inputs.

API estimates use 30 days of traffic. Always-on compute uses an approximate 730-hour month; prototype uses 160 hours. Custom retains the current hours for you to edit. Prices alone do not establish capacity for the same traffic. Compute billing assumes one continuous allocation; repeated starts may cost more.

Official API and hardware feeds refresh through a one-hour cache. If a feed fails, the result retains its recorded date and a warning.

GPU memory fit is a planning estimate. It does not establish serving throughput.

Owner-only NVIDIA assistant: disabled. The public calculators are available.

Deep links: ?project=planner, ?project=classifier, ?project=rag, and ?project=evaluation on this app's direct URL.