Lira supports both public and private model deployments. For enterprises that can't send conversation data to a third-party API, the same LLM that powers your agents runs entirely inside your own VPC or data center.
Enterprise deployments · dedicated onboarding · custom SLAs
A private LLM is a language model deployed on infrastructure a single organization controls — its own servers, VPC, or data center — rather than called over the public internet through a shared third-party API. No prompt, transcript or customer data leaves that environment. Lira supports open-weight models (via TensorRT-LLM and similar serving stacks) deployed this way, alongside your choice of public model providers when private hosting isn't required.
Serve the model from your own GPU infrastructure — on-premise or in your cloud VPC.
Prompts, transcripts and customer data never leave your network boundary.
Deploy the open-weight model of your choice — swap models without re-architecting your agents.
Inference runs on infrastructure sized for your call volume, not a shared multi-tenant queue.
Every inference call is logged on infrastructure you already control and audit.
Configure a public-model fallback for resilience without changing your default private deployment.
Meet data-residency and regulatory requirements that prohibit sending customer data to external APIs.
Keep patient conversations within infrastructure covered by your existing compliance program.
Satisfy sovereignty requirements that mandate the model run within a specific jurisdiction.
No — private LLM deployment is a joint effort. Our team helps size and configure the serving stack; your infrastructure team owns the environment it runs in.
Yes. Model choice is configured per agent, so you can run a private LLM for regulated flows and a public model elsewhere.
Depends on model size and call volume — our team sizes this with you during a technical scoping call rather than quoting a generic number.
Talk to our team about your infrastructure, compliance and scale requirements — we'll design the deployment around them.
Book a demo →