Six Layers of Redundancy for COB LED Video Walls in Control Rooms

What each layer protects, what it costs, and how the layers add up to 99.9% availability.

A video wall in a control room is not digital signage. When a pipeline SCADA alarm, a grid frequency excursion or a security incident appears on the wall at 3 a.m., operators act on what they see, and a dark screen is an operational incident in its own right. COB (chip-on-board) packaging has already removed a large share of the classic failure modes at pixel level: the LED chips are bonded directly to the board and sealed under a common epoxy layer, so there are no individual SMD packages with solder joints that crack under thermal cycling or a careless touch. But the electronics behind the pixels, meaning the power supplies, receiving cards, data cables and controllers, fail the way all electronics eventually do. Redundancy is the engineering discipline that makes sure a single failed part never becomes a dark wall.

The logic is the same as in a modern airliner: no single component is trusted with the whole mission. This article walks through the six practical layers of redundancy available in current LED control systems (using Novastar as the reference architecture, though Colorlight, Brompton and others offer equivalent mechanisms), weighs the pros and cons of each, shows how they combine in real projects, and explains what the popular tender phrase “99.9% system availability” actually obliges a supplier to deliver.

The signal chain in one minute

A typical system looks like this: a video source or processor feeds an LED controller (the “sending” side); the controller’s output ports send data over Cat6 cable, or over optical fiber with converters, to the first LED cabinet; inside each cabinet a receiving card decodes its slice of the image and drives the LED modules; cabinets are daisy-chained, so data passes from one receiving card to the next. In parallel, each cabinet has its own power supply units fed from the mains. Every element in that chain is a potential single point of failure. The six layers below eliminate them one by one.

Layer 1: Signal loop-back in the data chain

In a standard installation the data chain simply ends at the last cabinet. With loop-back, the output of the last receiving card is cabled back to a second, “backup” port on the controller. The chain becomes a ring, and data can enter it from both ends. If any cable or connector in the chain fails, the cards on the far side of the break are instantly fed from the other direction. Think of a ring road around a city: if one section is blocked, traffic still reaches every district by driving the other way around. One nuance deserves a straight answer, because it defines what this layer can and cannot do: loop-back protects against a broken cable or connector. If a receiving card itself fails, the loop keeps every other cabinet in the chain alive, fed from both ends, but the failed card’s own cabinet goes dark, because no data path can revive a card that no longer drives its modules.

Advantages: this is the cheapest layer of all, essentially one extra cable and a spare port per chain of cabinets, and it covers the most common field failure: a damaged cable or a loose RJ45 connector. Switchover is automatic and seamless, with no operator action and no visible glitch.

Disadvantages: the backup path consumes a controller port, which in practice means more output cards than the pixel count alone would require, or a larger controller model. Loop-back also has the blind spot described above: a failed receiving card still means one dark cabinet. And it does nothing against controller or power failures.

Verdict: the baseline. No control room wall should ever be quoted without at least this layer.

Layer 2: Dual receiving cards in each cabinet

This layer closes the blind spot of Layer 1. Each cabinet carries two receiving cards in hot standby, both wired into the main and backup data paths. If the primary card fails, the backup takes over within frames and the cabinet never goes dark.

Advantages: dual receiving cards are not an add-on to loop-back but its senior alternative: they cover everything the loop covers, a broken cable or connector anywhere in the chain, and add the one protection the loop cannot give, survival of a receiving-card failure. The backup cards form a second, fully independent data chain fed from duplicated controller outputs (Layer 3). For the same reason the two layers are an either/or choice: combining loop-back with dual receiving cards adds controller ports and cost without adding coverage. A wall runs either Layer 1, or Layers 2 and 3 together.

Disadvantages: it doubles the receiving-card count, the internal cabling and the output ports occupied on the controller, requires a cabinet HUB design that supports it, and addresses a comparatively rare failure mode: receiving cards fail far less often than cables or power supplies, so the cost per avoided incident is higher than for Layers 1 and 4.

Verdict: specify it where even one dark cabinet is unacceptable, such as on-air broadcast walls or dispatch centers where any single tile may carry critical telemetry.

Layer 3: Dual output ports on the LED controller

On modern chassis controllers, duplicated outputs come in two forms. The first is an output card with built-in backup ports. Novastar’s H series card with 16x RJ45 and 2x Fibre outputs, for example, carries sixteen RJ45 ports as primary outputs and two optical ports that mirror them completely: OPT1 duplicates the data of Ethernet ports 1–8, OPT2 duplicates ports 9–16. A single card therefore ships with a ready-made backup of its own outputs. The second form uses two identical output cards in the chassis, one carrying the primary signal and one the backup. In both cases the backup outputs feed the backup receiving cards of Layer 2, which is why these two layers always travel together: duplicated ports need somewhere to land.

Advantages: automatic, invisible failover configured once in software; protects against failure of the port electronics, of a whole output card, and of every cable on the primary path.

Disadvantages: the same wall now needs twice the output ports, which can mean an additional output card or a step up to a higher-class chassis. A controller specified to use every port at full capacity cannot be retrofitted with port redundancy, which is why redundancy must be decided before the controller is chosen, not after.

Layer 4: Dual power supplies in each cabinet

Statistically, power supply units are the most failure-prone electronic component in an LED wall: they run warm, they run 24/7, and their electrolytic capacitors age fastest. In a genuinely redundant configuration two PSUs work in parallel, each sized to carry the full cabinet alone, normally sharing the load at under 50% each. If one fails, the other takes the whole load without interruption, exactly like a twin-engine aircraft that is certified to fly on one engine. There is a valuable side effect: a PSU loaded at half its rating runs cooler, and cooler capacitors age slower, so redundancy also extends service life.

A warning from tender practice: some cabinets marketed as having “dual power supplies” merely split the cabinet into two halves, each fed by its own PSU with no failover. If one dies, half the cabinet goes dark. Specifications should explicitly require 1+1 redundancy with automatic switchover, not just “two PSUs.” And remember the upstream side: two redundant PSUs on the same circuit breaker still share a single point of failure. Dual mains feeds (A/B circuits, ideally from separate UPS) complete this layer. A side note for the sending end: if the project relies on a single LED controller, specify it with dual hot-swappable power supplies as well; chassis controllers such as the Novastar H series offer a redundant PSU as a factory option.

Disadvantages: added cost and heat, slightly lower PSU efficiency at partial load, and no protection against signal failures: a dead port or a cut data cable will darken a perfectly powered cabinet. The cost argument, however, cuts the other way: the premium for the second PSU is small compared with the price of a COB cabinet, which is why dual power supplies deserve to be a default requirement in every control room project, not an option.

Layer 5: Redundant fiber between equipment room and wall

In many control rooms the video processing racks sit in a separate equipment room, tens or hundreds of meters from the wall, for reasons of noise, heat, security and serviceability. Copper Cat6 is limited to about 100 meters and picks up electromagnetic interference along the way; single-mode fiber carries the same data up to 10 kilometers, immune to interference. Controllers with optical outputs make this elegant: on the H series output cards the optical ports are factory-defined mirrors of the Ethernet ports, so the fiber run is born redundant. At the wall end, a fiber converter such as the Novastar CVT10-S turns the optical signal back into RJ45 outputs for the cabinets. The CVT10-S itself is built for redundancy: it carries two 10-Gbit optical inputs with hot-swappable SFP modules, OPT1 as the working path and OPT2 as backup, and ten Gigabit Ethernet ports toward the wall — a main-plus-backup design that pairs naturally with the duplicated optical outputs of Layer 3.

Advantages: covers the longest and most exposed part of the signal path; allows main and backup fibers to be routed along physically separate paths, which is the whole point, because two fibers in the same tray protect against a transceiver failure but not against the excavator or the tray fire; lets the controllers live in a secured, air-conditioned room where they can be serviced without entering the control room; and on high-resolution walls a pair of fibers replaces dozens of long copper runs, simplifying containment and saving cable cost.

Disadvantages: the converters are additional active devices that can themselves fail, although a failed converter is a quick swap provided the customer keeps a spare unit on site, and in critical systems two converters are used, cross-connected to both fibers. Fiber also brings extra cost for SFP modules, splicing and testing.

Layer 6: Dual LED controllers

After Layers 1 to 5, one large single point of failure remains: the controller itself. Device-level backup removes it. Two identical controllers receive the same video signal, typically from a video processor with mirrored outputs. The primary controller feeds the main ends of the data chains, the backup feeds the backup ends, and the system switches over automatically if the primary stops transmitting. The same mechanism works across fiber: main controller on fiber A, backup on fiber B.

Advantages: covers complete controller failure, including its internal power supply, mainboard and firmware crash. It also enables zero-downtime maintenance: switch to the backup, update or replace the primary, switch back. For a 24/7 room this is often the argument that justifies the cost, since firmware updates otherwise require a maintenance blackout.

Disadvantages: doubles the controller budget; the device that splits the video feed becomes the new single point of failure and may itself need redundancy; and it demands configuration discipline. Every change to the screen configuration must be written to both controllers.

Summary of the six layers

Layer Protects against Cost impact When to specify
1. Signal loop-back Broken data cable or connector in the chain Minimal (extra cable, uses a backup port) The minimal level of protection, for every control room
2. Dual receiving cards Failure of a receiving card (one dark cabinet), plus everything Layer 1 covers Moderate (doubles receiving cards, controller output ports and cabling) Zero-tolerance walls: on-air broadcast, critical dispatch
3. Dual controller ports Port or output-card failure; hot-standby backup data path Moderate to high (more output ports: extra output cards or a higher-class chassis) Always together with Layer 2 (dual receiving cards)
4. Dual PSU (1+1) Power supply failure; extends PSU service life Moderate (second PSU per cabinet and for the controller) Always for 24/7 operation; insist on true failover
5. Redundant fiber + converters Damage on long runs between racks and wall Moderate (SFPs, converters, splicing) Remote equipment room, or high-resolution walls where fiber cuts the number of long cable runs
6. Dual controllers Complete controller failure; zero-downtime maintenance Doubles controller budget Mission-critical 24/7 rooms

 

Combining layers: three practical tiers

Redundancy is not all-or-nothing, and buying every layer for every project wastes money. In practice three tiers cover most control room requirements.

    • Tier 0, entry level: signal loop-back. The minimum acceptable for any control room: cable and connector failures, the most frequent ones, are covered for the price of a few patch cords and spare ports.

    • Tier 1, standard control room: port redundancy together with dual receiving cards (Layers 3 and 2, which replace loop-back and exceed it in coverage), and dual power supplies in the cabinets and in the LED controller. This covers cables, ports, cards and power, typically for a 10–15% premium on the display subsystem.

    • Tier 2, mission-critical 24/7 room: Tier 1 plus dual controllers, necessarily of identical type with the same number and type of output ports, plus redundant fiber where the racks are remote. Every active device between source and receiving card is now duplicated.

Each successive tier removes a progressively rarer failure mode at a progressively higher price, so the right question is never “how much redundancy exists?” but “which failure modes remain uncovered, and can we live with them?” One rule applies to all tiers: redundancy that has never been failover-tested is a hypothesis, not protection. Every layer must be demonstrated by deliberately failing the primary path during commissioning.

What “99.9% system availability” really means

Tender documents for control rooms routinely demand “99.9% availability,” and suppliers routinely promise it without either side calculating anything. The arithmetic is simple: for a system that must run 24/7, 99.9% allows about 8 hours 46 minutes of downtime per year. In other words, a single failure that takes one working day to diagnose and repair consumes the entire annual budget of a 99.9% commitment.

The number alone, however, is meaningless until the tender defines three things. First, what counts as “unavailable”: a full blackout, a single dark cabinet, more than a defined percentage of screen area, or even loss of redundancy while the image is still perfect? Suppliers and customers regularly discover at acceptance that they assumed different definitions. Second, whether agreed maintenance windows are excluded from the calculation. Third, how availability is measured and over what period, and what remedy applies if it is missed.

Meeting the requirement rests on the standard availability formula of reliability engineering: availability equals MTBF divided by the sum of MTBF and MTTR, that is, mean time between failures set against mean time to repair. The formula gives the designer two levers. Redundancy works on the first: with Tier 1 or 2 in place, most component failures cause zero visible downtime, and the repair becomes a scheduled swap instead of an emergency. The second, repair time, is at least as important and much cheaper to improve. A spare parts package on site (a common benchmark is around 5% of modules plus at least one of each electronic component type), front-serviceable COB modules that swap magnetically in minutes, remote monitoring that reports a failed PSU before anyone sees it, and staff trained to perform the swap: together these reduce MTTR from days to minutes. A wall with modest redundancy, good monitoring and spares in the cupboard will beat a heavily redundant wall whose only spare parts are four weeks away in a factory.

COB technology itself contributes quietly to the availability budget: with the LED chips sealed under a common protective layer, the pixel-level failure and repair events that plague fine-pitch SMD walls (knocked-off LEDs, corroded solder joints) largely disappear. The contrast with LCD video walls is even sharper: when an LCD panel fails, the entire panel must be replaced, and extracting a faulty panel from the middle of a wall is a challenging task for at least two people, typically with the wall out of service for the duration. A COB module, by comparison, swaps from the front in minutes, single-handed. The maintenance statistics therefore concentrate on the electronics, which is exactly where the redundancy layers stand guard.

Practical advice for installers

None of the points below requires expensive hardware, yet each one can significantly reduce the downtime of a video wall when a failure eventually happens, often turning an hours-long emergency into a swap of minutes.

    • Distribute the load evenly across output ports. Do not fill some ports to their pixel limit and leave others empty. Balanced, shorter chains keep bandwidth headroom for redundancy and higher refresh settings, reduce the number of cabinets affected between fault and failover, and make fault-finding faster.

    • Pull spare cables while the trays are open. At least one spare Cat6 run per cable route, and spare fiber cores on every optical route. Copper pulled during installation costs a few dollars; the same cable pulled after the ceiling is closed costs a service visit and a maintenance window.

    • Terminate through RJ45 patch panels, not direct crimped runs. With a patch panel at the rack and keystone couplers at the wall edge, replacing a suspect cable or re-routing a chain becomes a two-minute patch-cord change, and failover tests are plug-level operations instead of re-terminations.

    • Label both ends of every cable with controller, port and chain position, and leave an as-built cabling diagram and photos inside the rack door. The person troubleshooting at 3 a.m. will not be the person who installed the wall.

    • Test every failover at commissioning. Unplug each primary cable, switch off one PSU per cabinet type, power down the primary controller and the primary fiber, and witness the wall staying lit each time. Record the results in the acceptance protocol.

    • Keep pre-configured spares on site. Receiving cards flashed with the correct firmware and cabinet configuration file swap in minutes without a laptop, and a spare fiber converter turns a converter failure into a plug-level fix. Store configuration backups somewhere other than the control PC itself.

    • Separate the A and B power feeds onto different breakers and, where possible, different UPS, and balance the phases; redundant PSUs on one breaker are redundant in name only.

    • Enable monitoring with alerts. Temperature, voltage, card and link status should raise an email or SNMP alarm. Redundancy hides failures by design: a wall silently running on its backup path has already lost its protection, and nobody knows. Redundancy without monitoring is a parachute nobody ever inspects.
    •  

Conclusion

Every layer described here exists because somewhere, at some time, exactly that component failed on a wall that mattered. The good news is that none of the layers is exotic: loop-back, duplicated output ports and cards, 1+1 power, redundant fiber with converters like the CVT10-S, and dual controllers are catalogue features of mainstream control systems, waiting to be specified. The discipline lies elsewhere: decide the tier that matches the room’s criticality before the controller count is fixed, define availability in numbers rather than adjectives, put spares and monitoring in the contract alongside the hardware, and prove every failover with a pulled plug at commissioning. Do that, and 99.9% stops being a tender formula and becomes what it should be: roughly nine hours per year of theoretical downtime that, in a well-designed COB wall, operators will never actually see.

 

 

About the Author

This article was prepared by Michael Nevzorov, BDM of ACTM Visual B.V., a specialist in display solutions for mission-critical and control room applications with over 20 years of experience. ACTM Visual B.V. provides independent technology assessment, system specification, and tender support for EPC contractors, system integrators, and end-users across the utility, industrial, and transport sectors.