Ask a network team what keeps them up at night and they will name the visible things: the core routers, the internet edge, the customer-facing platforms. Almost nobody names DNS or DHCP. Yet these are the services that, when they fail, take everything else down with them, and they are usually the least looked-after infrastructure a Caribbean operator runs.
The reason is simple. DNS and DHCP are reliable enough that they disappear. A resolver set up five years ago keeps answering. A DHCP scope keeps handing out addresses. Nobody has a reason to open the configuration, so nobody does. The risk builds quietly, in three places.
The hardware ages out of support
The box carrying your resolvers reaches end of life. There is no patch path, no spares on the shelf, and no replacement plan agreed, because nothing has gone wrong yet. The first time this matters is the day the box fails, and by then the lead time on a replacement is measured in weeks.
The configuration drifts out of memory
Zone files, forwarders and scopes accumulate changes made under pressure, none of them documented. The person who understood how it all fits together has moved on. What remains is a working system that nobody can safely touch, which is a different problem from a broken one, and often a worse one.
The licence cost scales the wrong way
Commercial DNS and DHCP platforms frequently charge per query or per seat. That means the busier and more successful your network becomes, the more you pay for services that should simply scale with you. Growth arrives with a bill attached.
None of this shows up on a status dashboard, because everything is green until it is not. The failure mode is not a slow degradation you can watch coming. It is an outage that starts somewhere invisible and surfaces everywhere else first, so your team spends the first hour looking in the wrong place.
What good looks like
Redundancy that has actually been tested. Geo-redundant resolvers and DHCP failover pairs, so the loss of a single node or site is an event nobody outside the network team notices. Redundancy that has never been failed over on purpose is a hope, not a control.
Open standards you can read and keep. Configuration in open formats, on open-source software, so your setup is portable and your team can audit it without a vendor in the room. Standards-based infrastructure is not a purity argument; it is what keeps you from being locked in to a renewal you cannot walk away from.
Visibility before the outage. Query rates, lease utilisation and service health surfaced into monitoring, so exhaustion and drift show up as a rising line on a graph, days before they show up as a support call.
The uncomfortable truth is that most operators cannot say, today, which of their core services would break first. Not because they are careless, but because the services that never fail are the ones nobody is asked to examine.
That is what a network audit is for. It is the lowest-commitment way to find out where you actually stand: a read-only review of what you are running, what is out of support, what is undocumented and what would break first, delivered as a written report with no obligation to buy anything else.