AI in operations
Pipi Jarvis — an artificial intelligence inside the NOC/SOC team
The system continuously analyses data from the monitoring platforms and the logs, assesses anomalies and incidents and — within our permission and a set of predefined rights — makes decisions and carries out the actions they allow.
The problem
There is more data than the person on shift can possibly read
Thousands of metrics, a stream of logs and dozens of directions — every minute. A human cannot walk through that in order, so in practice they wait for an alert, and an alert only catches what somebody thought of in advance. Slow degradation — loss creeping up by a percent over 24 hours in one direction — goes unnoticed until a customer calls.
- On a night shift the analysis effectively stops
- An alert only catches what was anticipated
- Correlation across systems is done by hand
- The same check is repeated by hand every day
What we deliver
Pipi Jarvis — an artificial intelligence inside the NOC/SOC team
Continuous analysis
Monitoring and log data are processed continuously — not only at the moment an alert fires.
Anomaly assessment
The current state is compared with the baseline, and a meaningful deviation is judged in context.
Automated reports
Daily and incident reports, each with a conclusion and a priority.
Action within its rights
Pre-agreed actions, and only those; everything else goes to a human.
Escalation into the team
In the same channel the engineer on shift already uses — no separate interface.
Audit trail
What it assessed, what it decided and what it executed — recorded and reviewable.
Architecture
Architecture
- 01SourcesMonitoring, logs, routing state
- Zabbix
- Graylog
- smokeping
- BGP
- 02BaselineWhat counts as normal, per direction and per hour
- 03AssessThe deviation judged in context, not just a threshold
- 04RightsWhat it may execute alone, and what it may not
- scoped rights
- human-in-the-loop
- 05ReportA conclusion and a priority, in the team channel
- Telegram
- webhooks
- 06ReviewEvery action is reviewed and the boundaries tightened
- audit log
Operator consoles
Operator consoles
These exact systems run on the group's own infrastructure — the screenshots are processed before publication.
Global BGP Report — automatic, every day
From four international vantage points — Sofia, the USA, Amsterdam, India — every one of our prefixes is measured for reachability, loss and latency. No human assembles this: the system compares a 3-hour window against a 24-hour baseline, picks out the direction that is degrading and assigns the priority itself — in this frame P1 is India, where RTT passes 190 ms.
The team member that does not sleep
Pipi Jarvis sits in the same channel as the engineer on shift — it posts its reports and warnings there, answers questions and acts within the rights it was given. There is no separate interface: it lives where the team already works.
Capabilities
Capabilities
Every item is marked: verified production experience, or engineering capability.
Global BGP Report
ProvenA daily automated analysis from four international vantage points: reachability, packet loss and latency per prefix, a 3-hour window against a 24-hour baseline, with the degrading direction singled out and a priority assigned.
- BGP
- smokeping
- multi-region
Action only within granted rights
ProvenThe agent has an explicit list of permitted actions. Outside that list it does not act — it forms a recommendation and hands it to a human. Accountability stays with the engineer.
- scoped rights
- human-in-the-loop
ChatOps in the team channel
ProvenReports, warnings and answers where the team already works.
- Telegram
- bot API
Bringing the data sources together
ProvenMetrics, logs and routing state in a single assessment — the correlation nobody has time to do by hand.
- Zabbix
- Graylog
- BGP
First-line incident triage
CapabilityA proposed severity and a list of likely causes, ready before the engineer opens a screen.
Deployment into a customer environment
CapabilityThe same approach applied to your monitoring — with the boundaries of its rights defined by you.
Process
Process
- 01
Sources
Which monitoring and which logs it reads.
- 02
Baseline
What counts as normal, per direction and per time of day.
- 03
Rights
What it may execute on its own, and what it may not.
- 04
Reporting
Who receives the conclusion, in what form and how often.
- 05
Review
Every automated action is reviewed and the boundaries are tightened.
Technology stack
Technology stack
- Sources
- ZabbixLibreNMSGraylogsmokepingBGP
- Delivery
- Telegramwebhookscron
- Boundaries
- scoped rightsaudit loghuman-in-the-loop
Engagement model
Engagement model
Retainer
Monthly engineering hours — specialist access on demand.
Co-managed
NetWizard and your in-house team together, with split responsibility.
Fully managed
We own the agreed operational scope end to end.
Critical operations
24/7 monitoring, response and severity-based escalation.
Use cases
Use cases
Slow degradation
Loss on one direction climbs by a percent a day. It has not crossed a threshold yet — but it is already at the top of the report.
The night shift
At three in the morning the analysis does not stop; a human is woken only when that is genuinely warranted.
The daily routine
The same recurring set of checks is automated, leaving the engineer the work that is actually engineering.
FAQ
FAQ
Does it act on its own?
Only within boundaries we defined in advance. Outside that list it does not execute anything — it forms a recommendation and hands it over. Every automated action is logged.
Does it replace engineers?
No. It reads what a human physically cannot keep up with and hands the engineer a formed picture. The decision, and the responsibility, stay with the human.
What data does it see?
What it is explicitly given: monitoring metrics, logs and routing state. In a customer deployment that list is approved by you.