Corporate LLM Stack
Running business AI through a public service means quotes, contracts and customer data leave the building. Your own stack answers the same questions on the same documents — except the answer stays where the question was asked.
Your own language models under your control — on hardware in-house or on rented GPU capacity inside the EU. Inference, RAG and voice run GDPR-compliant, without a detour through somebody else’s consumer API.
- ›On-prem inference on RTX 5090 with vLLM, Ollama, llama.cpp
- ›RAG over your documents with Qdrant & OpenWebUI
- ›Voice agents for the phone with Pipecat & Asterisk
- ›Automation & agent workflows, integrated into your systems
Businesses with documents that must not go outside: quotes, contracts, personnel files, engineering drawings. And anyone who has noticed that a chatbot without access to your own data is not much help.
What we run ourselves
We run language models on our own graphics cards in continuous operation, not as a test setup. What we set up for you has been running on our end for months.
- Our own model servers
- Multiple machines with the latest graphics cards that provide open models around the clock. No tokens ever leave our premises.
- One access point for everything
- A gateway in front of the system bundles our own and third-party models behind a single interface. Switching providers is therefore just a matter of changing a configuration line—not a project.
- Queries on Your Own Documents
- Manuals, contracts, and files become searchable without being uploaded to a third-party provider.
- Automatically reading documents
- Invoices and delivery notes are analyzed and transferred to the inventory management system—with a verification step, not blindly.
- Email Classification
- A model evaluates incoming messages at the mail gateway before they reach the inbox.
How a project runs
The most common mistake is to start with the hardware. We start with the task—and it often turns out that a smaller model is sufficient and the cost estimate looks completely different.
- 1Refining the use case
What task, what data volume, what fault tolerance. We calculate the costs for the use case using both rented and in-house computing power.
1 Meeting - 2Measure, Don’t Guess
We test several models on your real data. The top-of-the-line model wins less often than the advertising suggests—and almost never in bulk processing.
1–2 weeks - 3Set up
Model server, gateway, access control, logging, integration with your applications. If needed, with a fallback to the EU cloud for peak loads.
2–5 weeks - 4Operation
Monitoring, model updates for new versions, cost control. Models become obsolete every six to eight weeks—the system must be able to handle this.
Ongoing
What it costs
The honest answer depends on one number: how much data you process per month. Below that threshold, renting is worth it; above it, your own hardware is worth it—we run the numbers for both options before making a purchase.
- Rented
- No upfront investment, pay-as-you-go billing, ready to go immediately. Pays off with fluctuating or low loads—and for data that’s allowed to leave the premises.
- Own Hardware
- One-time purchase plus electricity. Pays off starting at a constant base load, and is the only option where personal data remains securely on-premises.
- Hybrid
- Base load on your own hardware, with peak loads and special cases outsourced. In practice, this is the most common approach—and the reason why the gateway belongs at the front end.
What you get
- ✓A data-driven recommendation instead of a theoretical opinion
- ✓Operation on your own hardware, in the EU cloud, or a hybrid setup
- ✓A gateway that makes switching providers as simple as changing a configuration line
- ✓Logging that shows who requested what—a requirement under the EU AI Act
- ✓A cost calculation that includes operations, not just the initial purchase
Frequently asked
- Do we need our own GPUs?
- Not necessarily. One RTX 5090 carries a 30-billion-parameter model and often pays for itself against running API costs inside a year. Where buying hardware does not fit, the same stack runs on rented GPU capacity inside the EU.
- Are open models as good as the big providers?
- For questions about your own documents, usually yes — retrieval quality matters more there than model size. For open-ended creative work the large providers are ahead. Where an external model is the better choice, we say so.
- What is RAG and do we need it?
- Retrieval Augmented Generation means the model is handed the relevant passages from your documents before it answers. Without it a language model guesses; with it, it can cite the source. That is the difference between useful and dangerous.
- How does this stand under GDPR?
- With inference on your own hardware there is neither processing on your behalf nor a third-country transfer — the data never leaves the network. With rented EU capacity it stays an ordinary processing agreement inside the EU.
senn-tech