LatentStack/LatentGateway

LatentGateway

The right model access. For every team.

Configure supported endpoints, model tiers and team permissions. Keep restricted work on private models and allow approved third-party models for other workloads.

LatentStack is a modular AI software stack. This component is available on its own.
How it works
01Private and approved external models
02Team access and model tiers
03Permissions distinct from task routing
Illustrative model-access configuration
One request at the gatewayIllustrative · not live telemetry
guardrail L3 · PII scan · 1 email masked → forward clean
48,476Requests
9,154,504Tokens
$414.16Cost
1,308PII masked
Illustrative figures from a sample period, shown to explain what the gateway records per request: each request is priced, logged and attributed by user, project, model and provider. Not live telemetry, and not representative of an enterprise deployment's scale.

Supported model identifiers

Private · stays inside
on-prem/llama-3.1-70bon-prem/qwen2.5-72b
Hosted · crosses the boundary
anthropic/claude-sonnet-4-5anthropic/claude-haiku-4-5openai/gpt-4oopenai/o3-minigoogle/gemini-2.5-pro

Selecting a hosted provider sends the permitted inputs to that provider, under that provider's terms. Each identifier carries a verification date in the versioned compatibility list.

Works with
Models
Private self-hosted endpoints and approved third-party endpoints exposed over OpenAI-compatible APIs.
Agents
LatentCode, and supported MCP-compatible third-party coding agents configured to route through the gateway.
Identity
Team and role mapping from your existing directory, for access and model tiers.
What it does not do
It does not host or serve models. It governs access to the endpoints you provide or permit.
It does not remove the terms of an external provider. Permitted inputs sent to that provider are handled under that provider’s terms.
It does not route work by permission alone. Team permissions are distinct from task routing.
It does not screen every kind of personal data. L1 pattern and checksum rules can let some through, such as names in free text.
Deployment note
Available in all three deployment modes. In an air-gapped configuration there are no external model calls, so the gateway governs only locally available models.
Questions

Can we use both private and third-party models?

Yes, where your environment and policies permit. Configure supported private models for restricted work and approved third-party models for other workloads. Air-gapped deployments use models available inside the isolated environment.

What is screened before routing?

L1 pattern and checksum rules screen each input. Admins choose to block, mask or allow. Some personal data, such as names in free text, can pass.

Does our code leave our network?

It depends on the configuration you choose. Work restricted to private models has no external fallback. Work you permit to use approved third-party models sends the permitted inputs to that provider under that provider’s terms.

Evaluate LatentGateway on one of your repositories.