On-Premise AI As A Service: Your Models, In Your Building
We just launched a new service: ia.urbanodx.com, on-premise AI as a managed service. Open models, fine-tuned on your company's data, running on hardware that sits inside your building, operated by us so your team does not have to become an AI department.
The site is in Spanish because the first market is Spanish SMEs in regulated sectors. This post explains the thinking in English, because the argument is not Spanish at all.
Renting Intelligence Vs Owning It
Most companies today rent their intelligence. Every AI feature in the business runs through a metered API: usage grows every month, and the price per token is set by someone else. That arrangement has three structural problems.
The bill only goes up. API spend is a curve with one direction. You pay rent, rising rent, for a capability that never becomes yours.
Your data works away from home. Every call to an external API sends business information to third-party servers. For banking, healthcare, legal, and industrial companies, that is a risk no vendor clause fully removes.
Regulation pushes the other way. GDPR and the EU AI Act point in the same direction: the less your data leaves your organization, the less you have to justify. With the model in your building, "where is my data?" has a one-sentence answer.
Owning flips each of these. The hardware is a fixed asset in your office, the marginal cost of a query is electricity, and nothing your company knows goes out the door.
The reason this is now a practical choice, not a research project, is that open model families fine-tuned on a company's own data already match the paid APIs on concrete, well-scoped business tasks. Fine-tuning earns its cost on behavior rather than facts, and the facts come from retrieval over your own documents: the split we described in RAG vs fine-tuning.
What The Service Actually Is
Three phases, and the third one repeats on its own:
1. Evaluation and setup. We measure your case: which tasks, which data, what you currently spend on APIs. If the numbers work, we install the equipment in your offices and run the initial fine-tuning on site. Your data does not cross the door.
2. Managed operation. We run the system: monitoring, updates, support, and capacity. Your team just notices it has an AI that knows the business and does not bill per token.
3. Periodic retraining. Your business changes and the model changes with it. Twice a year we retrain on your new data, remotely and on your own hardware wherever possible.
The hardware comes in four sizes that all run the same stack: Starter (a pair of NVIDIA DGX Spark units, desk-sized, for companies up to about 100 employees), Pro (a cluster of four), Business (an NVIDIA DGX Station GB300 tower), and Enterprise (an NVIDIA HGX B300 NVL8 system with a dedicated room design). Because the software layer is identical, growing means changing iron, not redoing the project.
That software layer is the part we obsess over: open models from the Llama, Qwen, Mistral, and DeepSeek families chosen per task, fine-tuning on your data, RAG over your internal documentation with citations, user roles and audit logs, a usage panel with periodic quality evaluations, and a vLLM inference server exposing an OpenAI-compatible API. For most tools, migration is changing the URL and the key: the response starts coming from your own server room instead of the cloud.
No Public Prices, On Purpose
The service page shows tiers, hardware, and software, and no prices. Every proposal is calculated against your actual case and your current API spend, in writing, before anything starts.
The evaluation can also end with "keep renting". Below a certain level of API spend there is no business case for buying hardware, and when that is the honest answer, we say it plainly and propose something smaller instead: a sizing report, a roadmap, or simply cutting your existing API bill. We would rather lose a project than sell a bad one.
Do The Arithmetic Yourself
Sizing an AI machine is arithmetic, not magic: bytes per parameter, an overhead margin, concurrent users. We published the whole calculation as an interactive VRAM calculator, together with a plain-language guide to what GPU a company actually needs. Both are in Spanish, but the numbers speak every language.
Who It Is For
Spanish SMEs first, and inside that, the sectors where data leaving the building is a board-level question: banking, healthcare, legal, and industry. If you are reading this in English and the renting-versus-owning question applies to your company too, contact us; the evaluation conversation works the same in any language.