Datacenter & private cloud
A private cloud where losing a node is an incident, not an outage
Fabric design, storage, HA and provisioning automation — on the same principles as the platform running in our own datacenter.
The problem
The private cloud is one rack that simply kept growing
Servers were added whenever they were needed; there is one storage box and everything depends on it; the network runs through two switches whose configurations no longer match. Nobody counts capacity, because N+1 only ever existed in a slide.
- Storage is the single point of failure
- Host versions have drifted apart
- Machines are built by hand, without templates
- A restore has never been run end to end
What we deliver
A private cloud where losing a node is an incident, not an outage
Fabric design
Leaf/spine or collapsed core, LACP, MTU, a separate management network and out-of-band access.
Compute and capacity
Host count with N+1, workload placement and a growth forecast.
Storage topology
ZFS/RAID levels, NFS/iSCSI, snapshots, replication and disk health monitoring.
HA and a failover test
Losing a node is actually performed and the result is written down.
Provisioning and templates
cloud-init, templates and the API — a hand-built machine becomes the exception.
Backup and a restore rehearsal
Policy, retention and a scheduled restore into a scratch environment.
Architecture
Architecture
- 01Physical
- Redundant power
- Cooling
- OOB
- 02NetworkLeaf-spine
- EVPN-VXLAN
- LACP
- MTU 9000
- 03ComputeVirtualization cluster
- KVM / Proxmox
- VMware
- HA
- 04Storage
- Ceph
- ZFS
- NVMe
- 05RecoveryTested restore
- Snapshots
- Off-site
- DR runbook
Operator consoles
Operator consoles
These exact systems run on the group's own infrastructure — the screenshots are processed before publication.
The virtualization cluster
A multi-node Proxmox cluster: KVM virtual machines and LXC containers, NFS/HA storage, a separate backup NAS and migrated ESXi hosts. This website runs on it too — in a container, on the same infrastructure we sell.
8 clusters from one interface
Proxmox Datacenter Manager brings independent Proxmox VE clusters into one place. In our own infrastructure it manages 8 of them, spread across separate physical locations — 3 of those outside Georgia. The frame also shows the scale: 211 nodes online, 1,480 virtual machines and 127 containers — one view, one procedure.
vSphere — the cluster, its hosts and the workloads on it
A vCenter with several ESXi hosts and the services placed on them: web, billing, a domain controller, backup agents. This is the environment where VMware and Proxmox coexist — part of it still on vSphere, part already migrated. One team operates both, under one backup policy.
Storage — ZFS pools and replication
ZFS storage on a RAIDZ2 topology — 12 disks in one vdev, roughly 33 TiB usable, ECC memory with most of it serving as ZFS cache. And the part that matters: disks with errors, zero; scrub and scan, clean. This is the layer underneath the virtualization cluster and the backups.
Capabilities
Capabilities
Every item is marked: verified production experience, or engineering capability.
An own datacenter in Georgia
ProvenEnterprise-grade hardware with redundant power and cooling — the group's hosting platform runs there.
- Cisco
- Juniper
- Arista
- Dell
- HP
A KVM platform with hourly billing
ProvenRAID 10 SSD storage on Xeon Gold, with a dedicated IPv4 and a 1 Gbps port per virtual server.
- KVM
- RAID 10 SSD
- Xeon Gold
Multi-cluster Proxmox
ProvenEight independent Proxmox VE clusters in separate physical locations managed from one interface — three of them outside Georgia.
- Proxmox VE
- Datacenter Manager
- PBS
VMware vSphere clusters
Proven- vSphere
- ESXi
- vCenter
ZFS storage and replication
Proven- TrueNAS
- ZFS
- NFS
Backup and tested restore
Proven- Proxmox Backup Server
- Cyber Backup
- Snapshots
EVPN-VXLAN fabric
Capability- EVPN
- VXLAN
- Arista
- Open vSwitch
Distributed storage (Ceph)
Capability- Ceph
- NVMe
- CRUSH
Provisioning automation
ProvenIn the group's hosting a server is created from the API, with no operator involved.
- API
- cloud-init
- Control panel
Process
Process
- 01
Audit
Hosts, storage, network, licences and workloads.
- 02
Design
Target fabric, capacity and storage topology.
- 03
Build
Nodes, network, storage and backup — documented.
- 04
Migrate
Workloads in stages, with a rollback plan.
- 05
Validate
Losing a node and restoring are performed, not assumed.
- 06
Operate
Monitoring, capacity and the patch cycle.
Technology stack
Technology stack
- Compute
- Proxmox VEKVMVMware vSphereLXC
- Storage
- TrueNASZFSRAID 10 SSDNFSiSCSICeph
- Network
- AristaLACPOpen vSwitchEVPN-VXLAN
- Backup
- Proxmox Backup ServerCyber BackupZFS snapshots
- Automation
- cloud-initAnsibleTerraformAPI
- Monitoring
- ZabbixPrometheusGrafanaLibreNMS
Engagement model
Engagement model
Project
A one-off scope: audit, migration or implementation with a fixed outcome.
Retainer
Monthly engineering hours — specialist access on demand.
Co-managed
NetWizard and your in-house team together, with split responsibility.
Fully managed
We own the agreed operational scope end to end.
Use cases
Use cases
From one host to a cluster
Everything runs on one server and rebooting it stops the company. We build a cluster with shared storage and HA, and move the workloads in stages.
A second site for recovery
The backup sits in the same building as production. We add a second site with replication and write down which service comes back in which order.
Building a hosting platform
Selling virtual servers with automated provisioning and billing — exactly what runs on the group's own platform.
FAQ
FAQ
Do we need Ceph?
Often not. On a three or four node cluster, ZFS with replication is simpler and cheaper; Ceph earns its place when node count and growth justify it. The choice follows the capacity plan.
Can we host with you instead?
Yes — the group has its own datacenter and hosting platform. We look at both options: building on your site, or hosting with us, to the same engineering standard.
What RPO and RTO do you promise?
We do not quote a number before the design. Targets follow workload criticality and are proven by a tested restore — a figure announced without a test is not a promise.