BGP Flap Mitigation and Fast Convergence on MikroTik Routers

BGP Flap Mitigation and Fast Convergence on MikroTik Routers

A practical approach to stable routing, faster failure detection, and resilient MikroTik BGP design.

BGP instability costs more than most network operators realize. Every time a route or session flaps, traffic can take an unexpected path, sessions can reset, and customers may experience packet loss or degraded service.

For Internet service providers, wireless Internet service providers, municipalities, and enterprise networks, the goal is not simply to make BGP react faster. The goal is to detect genuine failures quickly, avoid reacting to harmless or transient events, and ensure that the entire routing design converges in a predictable manner.

The Cost of BGP Instability

Consider a border router with a momentary upstream or transport failure. If the failure is not detected promptly, traffic can continue toward a broken path. If the route repeatedly disappears and returns, the network can experience multiple withdrawals, advertisements, path changes, and unnecessary control-plane work.

That can create:

  • Packet loss and session timeouts during path changes
  • Asymmetric routing and unexpected traffic shifts
  • Congestion on backup paths that were not intended to carry all traffic
  • Repeated alerts and additional troubleshooting workload
  • Customer-facing service degradation and potential service-level impact

The important distinction is this: a route flap is repeated instability that must be investigated and controlled, while slow convergence is the time required for the network to recognize a legitimate topology change and install a replacement path. These are related problems, but they do not have the same solution.

Understanding BGP Convergence on MikroTik RouterOS

BGP convergence is the time required for the network to recognize a change, select a new best path, update the routing table, and begin forwarding traffic through the replacement path.

A typical convergence sequence includes:

  1. A physical, Layer 2, Layer 3, or forwarding-path failure occurs.
  2. The router detects the failure through interface state, routing-protocol timers, or Bidirectional Forwarding Detection.
  3. The BGP session or affected route is withdrawn or recalculated.
  4. RouterOS selects and installs the best remaining path in the RIB and FIB.
  5. Traffic begins using the new forwarding path.

Fast convergence requires every stage to work correctly. Aggressive BGP timers cannot compensate for a bad fiber, failing optic, duplex issue, unstable wireless link, overloaded router, or a design with no healthy alternate path.

Use BFD for Fast, Intentional Failure Detection

Bidirectional Forwarding Detection (BFD) is often the right tool when a BGP peer needs to recognize a real forwarding-path failure much faster than ordinary BGP hold timers allow.

BFD operates independently of BGP and can detect loss of bidirectional reachability in sub-second intervals. When BFD is correctly enabled and permitted for a BGP peer, RouterOS can use the BFD session state to take down the BGP session quickly and begin reconvergence.

Important: Faster detection is not automatically better. BFD timers that are too aggressive for a wireless, congested, or jittery path can create false failures and cause the very flapping they were intended to prevent.

In RouterOS v7, the design process should include:

  • Allowing BFD for the intended interface or peer in /routing bfd configuration
  • Enabling use-bfd=yes on the applicable BGP connection
  • Confirming that the peer supports and is configured for BFD
  • Testing realistic failure scenarios before production rollout
  • Monitoring BFD session state and state changes after deployment
/routing bfd configuration
add interfaces=sfp-sfpplus1

/routing bgp connection
set [find where name="upstream-primary"] use-bfd=yes

/routing bfd session print

BGP Timers: Use Them Carefully

BGP keepalive and hold timers remain important, particularly where BFD is not appropriate or is not supported by the remote peer. Lower timers can reduce failure-detection time, but they also increase control-plane activity and sensitivity to packet loss.

A practical design usually follows these principles:

  • Use BFD for critical, stable peers where sub-second detection is truly needed.
  • Use conservative BGP hold timers for peers that cannot support BFD reliably.
  • Do not apply one aggressive timer value to every peer in the network.
  • Document why a timer was selected and what link characteristics were considered.
  • Validate negotiated hold times from the active BGP session, not only from the configuration.

The right values depend on router capacity, peer count, transport quality, path latency, jitter, and the operational impact of a false positive. A fiber-connected core peer can often tolerate more aggressive detection than a remote peer across a variable wireless transport path.

Mitigate Route Flaps at the Source

The best BGP flap mitigation is to fix the cause of the flap. Before changing timers or attempting to suppress advertisements, review the physical path, interface counters, logs, routing dependencies, and route-origination policy.

Common causes of repeated BGP instability include:

  • Optical signal problems, damaged fiber, failing optics, or interface errors
  • Unstable wireless links, interference, weather-related fading, or inadequate link margin
  • Incorrect BFD settings or BFD enabled on only one side
  • Route origination tied directly to an unstable IGP or connected route
  • CPU or memory pressure during full-table processing or major route changes
  • Incorrect route filters, policy changes, or accidental prefix withdrawals

Where a prefix must remain advertised even if an internal dynamic path temporarily changes, consider a carefully designed route-origination policy. RouterOS can create blackhole routes for configured BGP networks using the network-blackhole option. This is useful only when the design clearly prevents traffic from being incorrectly attracted to the router during an internal failure.

Design caution: A blackhole route can prevent a BGP advertisement from disappearing, but it can also intentionally discard traffic if no better route exists. It must be designed and tested as part of the overall failure-domain and backup-path strategy—not added casually as a flap workaround.

Traditional BGP route flap dampening is also not a universal solution. Overly broad dampening can suppress legitimate routes and extend customer impact. For most networks, correcting the root cause, using sensible failure detection, filtering routes tightly, and maintaining clean route-origination policy produces better results.

MPLS, VPLS, VXLAN, and EVPN Considerations

BGP instability can affect more than Internet routing. In service-provider and enterprise designs, BGP may also carry VPN, VPLS, EVPN, or overlay reachability information. A control-plane flap can therefore impact multiple customer services or tunnel endpoints at once.

When BGP is part of an MPLS, VPLS, VXLAN, or EVPN design, test the full service path—not merely the BGP session. Verify that:

  • The underlay IGP converges correctly and reaches all required next hops.
  • Labels, tunnel endpoints, and overlay routes recover as expected.
  • Redundant paths have adequate capacity for failover traffic.
  • Customer Layer 2 and Layer 3 services remain stable during a controlled failure test.
  • Failure detection and recovery times meet the actual service requirement.

Deploy Changes in Phases and Monitor the Results

Do not enable fast timers or BFD across every BGP session at once. Begin with a non-critical or secondary path, establish a baseline, and test intentional failures during a maintenance window.

A phased rollout should include:

  1. Document current BGP session state, timers, route counts, and normal failover behavior.
  2. Deploy the proposed settings on a limited test peer.
  3. Test link loss, peer restart, and route withdrawal scenarios.
  4. Watch for false BFD failures, BGP session churn, packet loss, and unexpected traffic shifts.
  5. Adjust only after reviewing measured behavior.
  6. Expand to additional peers once the results are consistent and documented.

Useful monitoring points include:

  • BGP session uptime and session state changes
  • BFD session state, packet counters, and state-change history
  • Interface errors, optical levels, wireless signal quality, and link events
  • CPU and memory usage during convergence events
  • Route counts, route changes, and prefix-filter logs
  • Customer-facing loss, latency, and alert volume before and after changes

The Bottom Line

BGP stability is not achieved by selecting the lowest possible timer values. It comes from reliable physical infrastructure, clear routing policy, correctly designed redundancy, measured failure detection, and disciplined testing.

Link Technologies, Inc. helps network operators build BGP designs that converge quickly when a real failure occurs while remaining stable during normal operation. Whether you are managing a few external peers or a large service-provider core, the right design can reduce outages, alert fatigue, and operational risk.

How Link Technologies, Inc. Can Help

Link Technologies, Inc. provides ISP-grade and enterprise-grade BGP engineering, MikroTik RouterOS support, routing design, convergence testing, monitoring, network troubleshooting, and infrastructure consulting.

Explore Our Services Contact Our Team

Link Technologies, Inc.

IT Infrastructure • Wireless • Fiber • Hosting • Cloud Services
shop.linktechs.net

Leave your comment

*
Only registered users can leave comments.