Runtime guides

Installation, Usage & Runtime

Technical guides for deployment teams: which models to run, how to govern them, and what hardware to buy — written for banks putting agentic AI on their own infrastructure.

Curated by Kailas Lovlekar · July 2026. AI models, hardware and pricing are evolving at a rapid pace — information in these articles is for broad reference only.

3-part series

Deploying agentic AI on-prem at a tier 1 bank

PhantomOps runs as a hybrid platform — cloud model APIs where that's acceptable, and fully local, air-gapped inference where a bank's data residency and regulatory posture demand it. This series covers the three questions every infrastructure and risk team asks before sign-off: which open-weight model, how do we govern where it came from, and what GPUs do we actually need to buy.

01
Choosing a local model for a BFSI deployment
Permissive-license models (Qwen, Mistral, Phi) vs. custom-license models (Llama, Gemma) vs. deep-reasoning MoE — how licensing, accuracy and footprint should actually drive the choice.
Read the guide →
02
Security & governance: Hugging Face vs. Ollama
What actually gets scanned before a model reaches you, the real risk of trust_remote_code, and why neither tool is "bank-ready" without an internal artifact registry in front of it.
Read the guide →
03
GPU & hardware sizing cheat sheet
A working VRAM formula, a tiered GPU table from a 3B model to a 70B cluster, indicative pricing, and cheaper alternatives for narrow, single-task agents.
Read the guide →

Sizing a deployment for your own environment?

We'll walk your infrastructure team through model choice, governance and hardware sizing for your specific workload.

Talk to us