On-Premise AI Compute for Malta’s Regulated iGaming and Finance Operators

The cloud is the right home for a great deal of AI work. It is not automatically the right home for regulated data, and at steady utilisation it is not automatically the cheaper one either.

If you run an iGaming operation, a fintech or a bank in Malta, you are being asked the same question from two directions at once. The business wants to put AI to work, on fraud and anti-money-laundering scoring, player or customer risk, churn, support automation, and increasingly on running models in-house rather than sending data to a third party. At the same time your compliance and risk functions are asking where that data physically sits, who can reach it, and what happens to continuity if a supplier has a bad day. Those two pressures meet at a single infrastructure decision: do the AI workloads and the sensitive data behind them run in someone else’s cloud, or on hardware you control. This article is about when the honest answer is on-premise, what that infrastructure actually looks like, and where the cloud still wins. Sirap has supplied, installed and supported Supermicro servers across Maltese operators for years, in DataCore software-defined storage roles, as storage servers, and for on-prem AI inference and training, so this is written from building these systems, not from a brochure.

Why would a regulated operator keep AI on-premise at all?

Because for regulated data the question is not only “what performs best”, it is “what can we prove and control”. Three things push a Maltese iGaming or financial operator towards on-premise or in-country infrastructure, and none of them is a fashion.

The first is data residency and control. Player data, payment data and the models trained on them carry obligations under the EU’s Digital Operational Resilience Act (DORA) for financial entities, the NIS2 directive, and the expectations of the MGA and the MFSA. Keeping the data and the inference that runs against it on hardware you own, in a known location, makes residency and access something you can demonstrate rather than something you have to take a supplier’s word for. For MGA licensees there is also a specific, separate obligation: a real-time replicated copy of regulatory data on an inspectable server in Malta, which we cover in our guide to the iGaming replication server requirement. The second is operational resilience. DORA in particular is built around the risk that a critical third party fails or is breached; infrastructure you run yourself, with your own continuity design, is one way to reduce concentration risk rather than add to it. The third is cost behaviour. Cloud GPU capacity is priced for flexibility, and that flexibility is worth paying for when your demand is spiky. When you are running inference more or less continuously, the meter never stops, and at steady high utilisation an owned cluster usually crosses below the cloud bill within its service life.

The honest version of the residency argument: on-premise does not make you compliant, and no server ever will. It gives you a location and an access boundary you control, which is one input to a DORA, NIS2 or MGA story your risk function still has to write. Sirap supplies the infrastructure; you and your advisers own the compliance position.

The distinction most buyers get wrong: training is not the workload that lives with you

It is easy to assume “AI infrastructure” means a giant training cluster, decide that is too much, and stop there. For most regulated operators the workload that actually needs to sit close to your data is not training, it is inference: the model running in production, scoring a transaction for fraud, flagging an account, answering a support query, every second of every day. Training a model is occasional, bursty and often fine to do in the cloud or on a smaller in-house box; inference against live regulated data is constant and latency-sensitive, and that is the part there is a real argument for keeping on-premise. Getting this distinction right changes the hardware conversation completely, because a steady inference workload is served by a sensible GPU server or two and fast storage, not by a hyperscale rack. It is the difference between a proportionate on-prem deployment and an overbuilt one.

What does the on-premise stack actually look like?

It comes down to three tiers, and a regulated operator rarely needs all of them at full scale. We build them on Supermicro because the platform range spans all three cleanly, from a 65W edge box to an eight-GPU server, on the same management and support footing.

1. GPU servers, to run the models

This is the compute that runs inference and, where you do it in-house, training. Supermicro’s GPU systems run from a 4U Universal GPU server that takes a flexible mix of accelerators, up to the X14 5U NVIDIA GPU server for heavier work. The point is to size the accelerator count to the inference load you actually have, with headroom, rather than to a benchmark. A pair of well-specified GPU servers covers a lot of real-world fraud, AML and risk inference.

2. Storage that keeps the GPUs fed, often software-defined

A GPU that is waiting on storage is the most expensive idle component in the building, so the storage tier matters as much as the compute. Here is the distinction worth drawing: a traditional SAN is not the only way to get resilient, high-performance shared storage. DataCore software-defined storage pools standard Supermicro servers and their NVMe into a redundant, high-availability block store, with synchronous mirroring between two nodes so a whole server can fail without taking storage down. For an operator that wants resilience and performance without being locked to a single proprietary array, running DataCore on Supermicro is a genuinely different shape of answer, and it is one of the configurations we run most. Where the requirement is pure capacity rather than block performance, the X13 4U storage server provides dense, economical bulk storage.

3. Edge systems, for branch and satellite sites

Not every workload sits in the main rack. Retail-facing arms, branch offices, regional sites and on-location operations often need a small, quiet, secure box that does local inference or acts as a secure gateway, without a data-centre around it. Supermicro’s April 2026 family of compact AMD EPYC 4005 systems is built precisely for this: the mini AS-E300-14GR in a 2.5-litre enclosure, the short-depth 1U AS-1116R-FN4 for branch back-end consolidation, and the slim-tower AS-3015TR-i4 that takes a single GPU for sites that need local acceleration. They draw as little as 65W, and they carry the security features that matter for distributed regulated kit: a TPM 2.0 hardware root of trust and AMD’s Secure Encrypted Virtualization, plus out-of-band management so you can run them in places with no IT staff.

Sovereign AI: marketing phrase, or something real?

“Sovereign AI” was everywhere at Computex 2026, including in Supermicro and AMD’s announcements around the AMD Helios rack-scale platform, and it is worth saying plainly what is real in it and what is not. The real part is the principle: your data and the models trained on it stay under your control and within a jurisdiction you choose, rather than dissolving into a global cloud. That principle maps exactly onto a Maltese regulated operator’s residency and resilience obligations, which is why the phrase is more than noise here. The unreal part, for almost everyone reading this, is the scale. You do not need a Helios rack or a hyperscale cluster to act on the sovereignty principle. You need a couple of GPU servers, resilient storage and a sound continuity design, run in Malta. The big-launch hardware is where the industry is heading and a useful signal of platform direction; it is not what a Maltese operator buys. We will tell you when you are about to overbuy.

What about power? GPU racks are not normal server loads

This is the part that gets underestimated. A populated GPU server draws far more, and far more variably, than the file and application servers most server rooms were sized for, and AI inference load swings hard as work arrives. That changes the power and cooling design, and it changes the continuity design. We pair AI compute with Eaton power protection sized for the real, peaky GPU load, online double-conversion so the rack never sees a transfer gap, with three-phase units such as the 93PS or 93T for a proper server room. For a regulated operator there is a second dimension: the UPS is itself a networked device on a critical system, so it falls within the same NIS2 and DORA security thinking as everything else. Eaton’s certified Network-M3 management card is built for that, and we cover it in our guide on whether your UPS is a cybersecurity risk. Power is not an accessory to an AI deployment; it is part of whether the thing stays up and stays compliant. Sirap has been an Eaton power partner since 1988, so we design the two together rather than bolting one onto the other.

When is the cloud still the right call?

Often, and we will say so. If your AI demand is spiky or experimental, if you are still finding out which models earn their keep, or if you need to burst to dozens of GPUs for a short training run and then stop, the cloud’s pay-for-what-you-use model is exactly right and buying hardware would be a mistake. Plenty of the best designs are hybrid: train and experiment in the cloud, run steady production inference on-premise next to the regulated data, and keep the option to burst. On-premise earns its place when utilisation is high and steady, when residency or resilience obligations make a controlled location worth real money, and when you have the power, space and support to run it properly. If you do not have all three, a smaller on-prem footprint plus cloud is usually the honest answer, and that is the recommendation we will give even though it sells less hardware.

Backed by SirapCare and a local team

Buying servers is the easy part; keeping a GPU and storage estate healthy in production is the work. Sirap specifies, supplies, installs and supports this infrastructure with engineers based in Malta, which for a regulated operator means a known local partner inside your continuity and incident plans rather than a distant support queue. We run Supermicro nodes behind DataCore storage and on-prem AI workloads for operators here already, so the configuration is one we maintain, not one we are trying for the first time on your site.

The outcome

Done properly, an on-premise AI deployment gives a regulated Maltese operator three things at once: the AI capability the business is asking for, a data location and access boundary the risk function can actually stand behind, and a cost curve that, at steady utilisation, beats renting the same compute indefinitely. Done badly, it is an overbuilt rack that draws too much power and sits half-idle. The difference is entirely in the sizing and the design, the GPU count matched to real inference load, storage that keeps those GPUs fed, power that holds, and an honest split between what belongs on-premise and what belongs in the cloud. That design work is the part worth getting right before any hardware is ordered.

Talk to us about on-premise AI infrastructure →

Featured Products in this article

HOW WE DO
Picture of Kurt Paris

Kurt Paris

With an MSc in Software Engineering, and over 15 years in IT Management, Kurt Paris leads technology strategy at Sirap. Zebra Technologies, Domino and Cisco-certified, he helps Maltese businesses build resilient storage & backup infrastructure, Machine Vision & AutoID Automation

Tags

What do you think?

Related articles

Contact us

Partner with Us for Comprehensive IT

We’re happy to answer any questions you may have and help you determine which of our services best fit your needs.

Your benefits:
What happens next?
1

We Schedule a call at your convenience 

2

We do a discovery and consulting meting 

3

We prepare a proposal 

Schedule a Free Consultation