Book a 20-min call

TECHNOLOGY SERVICE · 12

Managed Support & SRE

Ongoing operation of production systems: monitoring, on-call response, patching and the steady removal of whatever keeps breaking. For organisations without a night shift, and for teams whose engineers are spending their week on tickets instead of roadmap.

Service levels you can hold us to

Support begins with written definitions: what counts as a severity one, when the clock starts, what response and restoration times apply, and how that is reported each month. Vague promises of availability help nobody at two in the morning.

Coverage is agreed to match the system. Business hours in Gulf Standard Time suits many internal platforms, while customer-facing services usually justify round-the-clock on-call.

  • Severity definitions agreed with your business owners
  • Response and restoration targets per severity
  • Monthly report on incidents, service levels and outstanding risk

Alerts a human should act on

Most support handovers arrive with hundreds of alerts and no signal. Monitoring is rebuilt around symptoms users notice, expressed as error budgets and latency objectives, so a page means something is genuinely wrong. Everything else becomes a dashboard or a ticket.

Onboarding a new system takes two to four weeks: reading the code, documenting the runbooks, fixing the observability gaps and rehearsing a failure before we accept the pager.

Incidents end with a written review

Every significant incident produces a blameless review with a timeline, the contributing factors and actions with named owners and dates. Those actions are scheduled like any other work, because a review whose actions are never done is just paperwork.

Toil gets engineered away

A fixed share of each month goes to removing repeat work: the manual restart, the certificate nobody renewed, the job that fails every Tuesday. Support cost falls over time because the underlying causes are fixed, and that trend is visible in the monthly report rather than asserted.

What you get

  • Signed service level agreement with severity definitions and targets
  • Monitoring, logging and alerting tuned to user-visible symptoms
  • Runbooks for the failure modes we can foresee, tested in a drill
  • On-call rota with a documented escalation path to named people
  • Patch and dependency update schedule, applied and evidenced
  • Backup and restore verification on a stated cadence
  • Blameless incident reviews with tracked remedial actions
  • Monthly service report covering incidents, spend and risk

Typical outcomes

99.9%

Uptime commitment under a standard agreement

24/7

Coverage available for customer-facing systems

15 min

Target acknowledgement for severity one incidents

Stack we use

PrometheusGrafanaOpenTelemetryDatadogPagerDutyElasticKubernetesTerraform

Questions

Yes, after a two to four week onboarding that covers documentation, observability gaps and a rehearsed failure. Anything we consider unsafe to operate is raised before the agreement starts, not after the first outage.

A monthly fee based on the systems covered, the coverage window and the severity targets, with a defined allowance of improvement work included. Out-of-scope project work is quoted separately rather than absorbed quietly.

Yes for agreements with round-the-clock coverage. Business-hours agreements follow the Monday to Friday UAE working week, with holiday cover listed in the contract so there are no surprises around Eid.

Runbooks, dashboards and automation live in your repositories throughout, so a transition is a handover rather than a rebuild. Agreements run monthly after any initial term, and we will train your incoming team as part of the exit.

Next step

Start with a 20-minute call.

Tell us the roles you need filled, the system you need built, or both. You will speak to someone who has done the work, and leave the call with a route forward.