Deploying on your own hardware
Polaris on hardware you own needs no cloud attestation service, no cloud identity, and no outbound internet once the artifacts are in place. Read Deploying Polaris first for the choices common to every environment. This page covers what is specific to bare metal, including the air-gapped case.
On-premises deployments are delivered as engagements with our team, so you will not be doing this alone. This page is what we work through together, and it is written so your own operators can run it afterwards.
What this gives you, and what it does not
Worth being plain about the boundary before anything else, because it differs from the cloud deployments.
On your own hardware you get memory confidentiality for the workload against host software, since SEV-SNP encrypts guest memory with a key the hypervisor cannot read. You get integrity monitoring, because the sentinel watches syscalls, network destinations, kernel modules and the IMA log while the Policy Manager appraises what it sees against a baseline. And you get key gating, because once a layer is promoted from advisory to enforcing, a failed appraisal stops the workload being served.
What it does not give you is a defence against whoever owns the machine. For this release the host is inside the trust boundary. The person with root controls the key manager, the baseline and the artifacts the guests are launched from, so an operator who wants the keys can have them. Polaris on-premises is there to detect drift in your own estate and to gate keys on that, not to protect a workload from its own infrastructure owner. If your case has the workload owner and the host owner as different parties, tell us, because that needs a different design.
Before you start
You need an AMD EPYC server with SEV-SNP, Milan generation or later, and SEV-SNP enabled in firmware. On most vendors that is more than one setting, typically SEV, SEV-SNP and an SNP memory reservation, so consult your vendor's documentation rather than assuming a single toggle. Two guests is the minimum, one for the workload and one for the Policy Manager, so you need the ASIDs for both.
Confirm the host agrees before going further:
cat /sys/module/kvm_amd/parameters/sev_snp # expect Y
ls -l /dev/sev # must exist
dmesg | grep -i 'SEV-SNP' # the firmware version the PSP reports
If sev_snp reports N or the parameter is absent, the fault is in firmware or the kernel and nothing below will work.
Record three numbers from the machine that will run the guests:
grep -E '^cpu family|^model[[:space:]]|^stepping' /proc/cpuinfo | head -3
These go into the configuration and they are required. CPUID 1 EAX appears in every VMSA and every VMSA is covered by the launch measurement, so a set of artifacts is valid for one CPU model. Launching them on a host whose family, model or stepping differs is refused rather than silently accepted.
The reference platform is the RHEL 10 family. Validated on Rocky Linux 10.2 with libvirt-daemon, qemu-kvm, virt-install, guestfs-tools, edk2-ovmf, python3 and cpio installed. rpm2cpio comes from rpm itself, so rpm-build is not needed.
SELinux
If you run SELinux Enforcing, which is the RHEL default, the guests will not start until the staged firmware is labelled. qemu runs as svirt_t and the artifact directory is not a path the libvirt policy knows, so the firmware lands as var_lib_t and nothing permits that read. The failure looks like this and does not mention SELinux:
error: internal error: QEMU unexpectedly closed the monitor
qemu: could not load PC BIOS '<artifact_dir>/polaris-ovmf.fd'
Label it as content libvirt may read, which is what the tool will do for you in a later release:
semanage fcontext -a -t virt_content_t '<artifact_dir>/polaris-ovmf\.fd'
restorecon -v <artifact_dir>/polaris-ovmf.fd
The guest disks need nothing, because libvirt labels those itself when it starts a domain.
One more under Enforcing: the guests' serial console logs are written into the artifact directory, and virtlogd is not permitted to create files there. Writing to an already-open log keeps working, so this bites when a guest starts and when a log rotates, which is exactly when you want the console. Until the tool moves these under /var/log/libvirt, either label them too, with virt_log_t, or read the console with virsh console <domain> instead of the file.
If something does not start under Enforcing, the host has already recorded why. ausearch -m AVC -ts recent or grep 'avc: *denied' /var/log/audit/audit.log names the process and the path. In Permissive the same lines appear with permissive=1, which means the access was allowed but would have been blocked, so a Permissive host will tell you what Enforcing is going to break before you switch.
To ask the tool whether a host is ready before committing to anything:
./scripts/deploy.sh deploy --preflight --cloud onprem --config onprem.yaml
That checks what the bake needs on this machine and what a launch needs, and says which of the two this host can do. A build machine that only produces artifact sets for somewhere else will report that it can bake but not launch, which is expected and not a failure.
TEE options
AMD SEV-SNP, verified directly against AMD's published keys. There is no cloud attestation service in the path and no external attestation authority, which is the property that makes this configuration suitable for air-gapped and sovereign environments.
Recent AMD generations also support SVSM-based key isolation, which moves the guest's virtual TPM inside the trust boundary. Confidential GPUs are supported on-premises as they are in the clouds, with the caveat that NVIDIA's attestation service is the one outbound dependency Polaris cannot remove when they are in play.
How you deploy
Start from the shipped on-premises example configuration and read its comments, which are the reference for what each key changes. Five values have no default and the tool stops without them: an NTP server, the Policy Manager's address, and the three host CPU numbers above. There is no metadata service on bare metal, which is why the time source and the Policy Manager address have to be stated rather than discovered.
To see what a configuration produces without building anything:
./scripts/deploy.sh deploy --dry-run --cloud onprem --config onprem.yaml
Then bake the artifacts and launch:
./scripts/deploy.sh deploy --cloud onprem --config onprem.yaml \
--image-tag <tag> --workload-dir ./workload
./scripts/deploy.sh launch --cloud onprem --config onprem.yaml
On this path the tool acquires no container images. It expects your Docker daemon to already hold them and exports exactly those references into the guest image, so if you build elsewhere then save, copy and load them on this host first.
The bake prints two launch measurements, one per guest, which differ because the vCPU counts differ. Both are computed from the same staged firmware. Container images are not part of a launch measurement, which covers the initial memory image rather than the disk, so rebuilding an image leaves both measurements unchanged. Catching a changed container is the job of the runtime layers and the IMA allowlist.
To take the guests down without discarding the artifacts, use destroy with the same configuration.
Verification path
Read the verdict from the Policy Manager through its own attested TLS front, using the read-only token the deployment generated:
curl -sk -H "Authorization: Bearer $READONLY_TOKEN" https://<pm_ip>:443/api/v1/status
curl -sk -H "Authorization: Bearer $READONLY_TOKEN" https://<pm_ip>:443/api/v1/checks
See Verifying a deployment for what each layer means.
Two things to expect. The first cycles will raise findings, because the learners converge over roughly a day and a converged deployment reports a passing status with no alerts. And the Policy Manager's alert list is the observer, not the guest's console log. A console log shows only what the containers chose to write, so when what you are checking is that something did not happen, read the alert list.
Baselines ship with every layer advisory, which means findings are raised and no key is ever withheld. Promoting a layer to enforcing is a deliberate step, and one you should take only once that layer reports clean, because enforcing a layer whose allowlist is still empty is either inert or immediately fatal depending on the layer. On-premises:
./scripts/deploy.sh promote --cloud onprem --config onprem.yaml \
--layers l31 --mode enforce
--layers takes any of l2, l3, l31 and ima_boot, comma-separated. --mode is enforce, advise or off. Add --dry-run to see the current dials and what the change would be without sending it, which also prints what each layer has actually learned so far.
Once a layer is enforcing, a failed appraisal stops the workload being served. There is no cloud key service to revoke on your own hardware, so what you observe instead is the proxy refusing traffic and its readiness route reporting unavailable, which takes the workload out of service behind any load balancer in front of it. The timing is a budget of failed reports rather than a single one, so expect a couple of minutes between attestation stopping and traffic being refused, and recovery within seconds of the cause being resolved.
Key management
Key release integrates with your own key manager rather than a cloud service. For HashiCorp Vault, Polaris uses the transit engine only and needs no enterprise features.
Unsealing your Vault is yours and never ours. Use Shamir shares distributed to holders you choose, or PKCS#11 auto-unseal where you already run an HSM. Polaris never holds an unseal share, because a Polaris component that could unseal the key manager that gates Polaris keys would make the gate decorative.
Two behaviours to size your policy around. A token issued from a certificate login outlives that certificate, and keeps working until its own lifetime expires even after the certificate has expired or been revoked, so the token lifetime you choose is the window a revoked credential keeps working. And revoking a certificate stops the next login without touching tokens already issued, so plan for both actions rather than one.
Air-gapped deployments
Everything Polaris would otherwise fetch has a local alternative: the guest firmware package, the guest base image, the guest's own packages, the container images, and AMD's certificate and revocation endpoints, which are replaced by certificates supplied locally. Your own NTP server covers time. The generated on-premises baseline contains no cloud metadata destinations and no cloud identity fields.
Keep local copies of both pinned upstream files, the firmware package and the guest base image. Each is verified by digest, so a substituted file is refused rather than used, but a file that has been removed upstream cannot be fetched at all, and distributions do prune superseded versions. If you may need to reproduce a build months later, copy both into your own storage now and point the configuration at them. A build host with no egress needs this regardless.
Also keep a local cache of the guest's packages. Without one the build resolves them against the distribution archive, which needs egress and is not reproducible, so two builds weeks apart will produce different guest images.
Keeping the firmware current
Polaris does not build its own guest firmware. It takes your distribution's packaged EDK2 firmware, verifies both the package version and the digest of the firmware image inside it against values pinned in the tooling, and stages exactly those bytes. The digest is the contract, because the baseline's launch measurement was computed for those bytes, so a mismatch is refused rather than shipped.
So when EDK2 has a vulnerability, the path is that your distribution ships a patched package, Polaris updates the pinned version and digest, and you rebuild. The firmware bytes change, so both launch measurements change and so does the Policy Manager's identity pin, which means a new set of artifacts rather than a package update inside a running guest. In an air-gapped site that is something delivered rather than downloaded. Your firmware therefore moves on your distribution's schedule for the patch plus a Polaris release for the pin, and it does not wait on us to build firmware.
On a Rocky 10.2 host you already have the firmware and can confirm you have the right bytes against the copy your own package installed:
sha256sum /usr/share/edk2/ovmf/OVMF.amdsev.fd
Notes
There is no SSH into the guests. They are measured and carry no authorized keys, so their serial consoles in the artifact directory are the way in.
If the guests boot, load their images, start every container and then fail attestation, check the platform setting. Bare-metal SEV-SNP has its own value, and a cloud platform value on bare metal produces evidence claiming to come from that cloud.
If every guest fails the launch measurement check, the artifacts were built for different firmware, a different vCPU count or a different host CPU. All three are covered by the measurement.
If the workload's proxy refuses the Policy Manager over a measurement mismatch, the Policy Manager identity pin in the artifacts is from an earlier build. A pin is only valid for the firmware and guest sizing it was computed from.