Monitoring & observability
Monitoring that answers "why", not just "what"
We build the monitoring platform from scratch or fix the one you have: metrics, streaming telemetry and flow analytics in one picture, with dashboards and alerts you can trust.
The problem
Plenty of dashboards, no answers
The usual picture: dozens of graphs, hundreds of alerts, and still no answer to which link saturated or which customer was affected. The cause is almost always the same — the collection layer is built at the wrong level.
- SNMP polling interval is slower than the event
- Flow data is not collected at all
- Alerts fire on thresholds, not on symptoms
- Dashboards match nobody's actual workflow
What we deliver
Monitoring that answers "why", not just "what"
Collection architecture
SNMP, gNMI and flow — each at the layer where it genuinely works.
Metrics backend
Retention, cardinality and scaling calculated up front.
Role-based dashboards
NOC, engineer and management — three distinct views.
Alerting rules
Symptom-based alerts with escalation routing.
Synthetic checks
Simulating the user path — HTTP, DNS, voice, stream.
Documentation
What is measured, why, and who responds.
Architecture
Architecture
- 01CollectPolling + streaming + flow
- SNMP
- gNMI
- NetFlow/IPFIX
- 02Process
- Telegraf
- Akvorado
- Vector
- 03StoreTime series and flow
- Prometheus
- InfluxDB
- ClickHouse
- 04Visualise
- Grafana
- Zabbix
- 05AlertRouting and escalation
- Alertmanager
- Telegram
- PagerDuty
Operator consoles
Operator consoles
These exact systems run on the group's own infrastructure — the screenshots are processed before publication.
Bandwidth and peering in real time
Every upstream and every peering session is measured separately, with global and local traffic split apart. Capacity and peering decisions come from this view — not from a hunch.
Where the traffic actually comes from — flow analytics
Flow records grouped into ASN categories: CDN/OTT, Georgian ISPs, the group's own networks and international transit. This is the picture that decides which cache to bring inside the network, where peering pays for itself and what stays on paid transit. It is the one screen published legible — everything in it is an aggregate share, with no customer and no address anywhere.
Who is actually sending the traffic — per ASN
The same flow data, this time with names on it: which CDN and which platform sends the most, where it lands and which router carries it. This list is what turns a peering or cache-placement conversation into an argument with evidence — Akamai, Fastly, Meta, Google. The volumes are blurred; the names and the proportions are not.
One map — MikroTik, Ubiquiti and Cambium together
A small or mid-sized operator's network is rarely single-vendor: MikroTik in the core, Ubiquiti on the wireless links, Cambium at the access layer. This map carries, per link, what operations actually runs on — wireless clients on the sector, CCQ, noise floor, airMAX quality and capacity, distance, frequency, Rx/Tx and uptime. The tower power nodes sit on the same map: battery voltage is a metric like any other here, and it explains half of the night-time outages.
Network inventory and health
Over 130 devices and thousands of ports in one system: an availability map, alert history and the top errored interfaces — including the GPON/EPON access layer, where a fault shows up on the port before the subscriber notices it.
One circuit, one graph — QoS in real time
This is what one business circuit looks like in Zabbix: inbound and outbound traffic, peaks, the real use of the committed rate, and history. It is the graph that answers "is the link actually enough" — and the customer can see that answer too, not only us.
Capabilities
Capabilities
Every item is marked: verified production experience, or engineering capability.
Site power and batteries
ProvenBattery voltage, the mains feed and temperature on the same map as the traffic. A sagging voltage is visible before the outage — and it explains the share of night-time incidents that otherwise stays unexplained.
- SNMP
- Battery voltage
- Temperature
SNMP polling at scale
ProvenPolling stays the baseline mechanism across most of the fleet — counters, availability and interface state.
- SNMP
- Zabbix
- LibreNMS
gNMI streaming telemetry
CapabilityStreaming telemetry replaces polling only where the sampling rate demands it — elsewhere the two complement each other.
- gNMI
- gnmic
- Telegraf
Flow analytics
CapabilityNetFlow/IPFIX/sFlow on ClickHouse for peering and traffic analysis.
- Akvorado
- ClickHouse
- IPFIX
Grafana dashboards
Proven- Grafana
- Prometheus
- InfluxDB
Zabbix infrastructure monitoring
Proven- Zabbix
- SNMP traps
- Agents
Device inventory and port monitoring
ProvenThe whole device fleet and its ports in one system — with autodiscovery, an availability map and ranked error counters, down to the GPON/EPON access layer.
- LibreNMS
- SNMP
- GPON/EPON
SLO and error budgets
CapabilityCapacity forecasting
CapabilityGrowth trends and saturation forecasting for links and storage.
Technology stack
Technology stack
- Metrics
- PrometheusInfluxDBZabbixLibreNMSVictoriaMetrics
- Telemetry
- gNMIIOS XR telemetryTelegraf
- Flow
- AkvoradoClickHouseNetFlow v9IPFIXsFlow
- Visualisation
- GrafanaZabbix UI
- Alerting
- AlertmanagerGrafana AlertsTelegram
Engagement model
Engagement model
Project
A one-off scope: audit, migration or implementation with a fixed outcome.
Retainer
Monthly engineering hours — specialist access on demand.
Co-managed
NetWizard and your in-house team together, with split responsibility.
FAQ
FAQ
Zabbix or Prometheus?
Both, for different jobs. Zabbix is strong for appliance and agent monitoring; Prometheus for dynamic, containerised environments. They frequently coexist under one Grafana.
How long is data retained?
Retention is designed against budget and requirement — high resolution short-term, aggregated data long-term.