Runtime guides
Installation, Usage & Runtime
Technical guides for deployment teams: which models to run, how to govern them, and what hardware to buy — written for banks putting agentic AI on their own infrastructure.
Curated by Kailas Lovlekar · July 2026. AI models, hardware and pricing are evolving at a rapid pace — information in these articles is for broad reference only.
3-part series
Deploying agentic AI on-prem at a tier 1 bank
PhantomOps runs as a hybrid platform — cloud model APIs where that's acceptable, and fully local, air-gapped inference where a bank's data residency and regulatory posture demand it. This series covers the three questions every infrastructure and risk team asks before sign-off: which open-weight model, how do we govern where it came from, and what GPUs do we actually need to buy.
01
Choosing a local model for a BFSI deployment
Permissive-license models (Qwen, Mistral, Phi) vs. custom-license models (Llama, Gemma) vs. deep-reasoning MoE — how licensing, accuracy and footprint should actually drive the choice.
Read the guide →
02
Security & governance: Hugging Face vs. Ollama
What actually gets scanned before a model reaches you, the real risk of
Read the guide →
trust_remote_code, and why neither tool is "bank-ready" without an internal artifact registry in front of it.03
GPU & hardware sizing cheat sheet
A working VRAM formula, a tiered GPU table from a 3B model to a 70B cluster, indicative pricing, and cheaper alternatives for narrow, single-task agents.
Read the guide →