How to Troubleshoot Common Issues in Waste Heat Recovery Boilers?
Waste heat recovery boilers are not forgiving when something goes wrong quietly. A fouled convective tube bank, a drifting drum level, a blowdown valve that hasn’t been exercised in six months — any of these can drag thermal efficiency from the 85–90% range down toward 70% before the control room notices a trend. Stack temperature creeping up 30°C over a quarter looks like a calibration drift until you’re replacing a superheater section or explaining an unplanned outage to a customer whose process line just went cold.
The most common issues in waste heat recovery boilers — fouling and scaling on heat transfer surfaces, drum level instability, tube corrosion from flue gas condensation, sootblower failures, and refractory degradation — are nearly always detectable early through stack temperature trends, steam quality checks, and differential pressure readings across tube banks. Catching them at the symptom stage costs a fraction of what reactive repair does.
What makes WHRB troubleshooting genuinely tricky is that the fault signatures overlap. A rising stack temperature could mean ash fouling on the convective passes, or it could mean you’ve lost sootblower coverage on one zone, or the flue gas flow has shifted because something upstream in the process changed. Same symptom, three different root causes, three different fixes — and the wrong call wastes time and can make the problem worse. The sections below work through each failure mode systematically, with enough operational context to tell them apart.
Diagnosing Abnormal Steam Output Drop and Low Boiler Efficiency
A drop in steam output is the single most-reported complaint from WHRB operators, and the frustrating part is that it can originate from at least half a dozen different places — some completely outside the boiler itself. Before you start chasing tube fouling or scaling, you need a baseline to measure against.
Establishing Your Performance Baseline
Know your design numbers cold: rated steam flow in t/h, flue gas inlet temperature, flue gas outlet (stack) temperature, and pinch point temperature difference. For most waste heat recovery boilers operating on industrial exhaust — cement kilns, glass furnaces, steel reheating furnaces — a healthy pinch point sits somewhere between 15°C and 30°C. Tighter than 15°C usually means you’re over-surfaced or running below design load; wider than 30°C is your first signal that something is degrading heat transfer.
Log these values every shift, not just during commissioning. A single week of trending data is worth more than any one-off spot check.
Stack Temperature as Your Primary Efficiency Sensor
Stack temperature is the cheapest and most reliable indicator of boiler-side efficiency loss. The rule of thumb: every 10°C rise in stack temperature above design baseline corresponds to roughly 0.5–1% efficiency loss, though the exact figure depends on flue gas composition and moisture content. On a boiler recovering heat from a cement preheater at around 350–450°C inlet, that 10°C stack creep can translate to meaningful lost steam tonnage over a month.
Set up a continuous thermocouple or RTD at the stack, tie it to your DCS or even a simple datalogger, and establish a control band — say, design stack temperature ±8°C. When you drift outside that band for more than two consecutive shifts, trigger a formal diagnostic.
A 10°C rise in stack temperature above design baseline typically indicates a 0.5–1% thermal efficiency loss in waste heat recovery boilers.True
This relationship is well-established in boiler engineering and follows from the flue gas heat balance: higher exit temperature means more enthalpy leaving the system unrecovered. The exact loss depends on flue gas mass flow, specific heat, and moisture content, so the 0.5–1% range is realistic for typical industrial WHRB conditions.
Separating Process-Side from Boiler-Side Causes
This step matters enormously and gets skipped too often. If the kiln upstream cuts production by 15%, flue gas flow drops, inlet temperature drops, and your steam output naturally falls — that’s not a boiler problem. Before you schedule any inspection, pull the process data: kiln throughput, combustion air flow, and upstream exhaust temperature over the same trending period.
If steam output is falling but process conditions are stable or even improving, you’re looking at a boiler-side cause. That narrows your list to tube fouling on the gas side, waterside scaling, or air in-leakage diluting the flue gas.
Tube Fouling Assessment — Step by Step
Start with numbers. Calculate the overall heat transfer coefficient (U-value) from your current operating data using the standard heat exchanger equation, then compare against the design U-value from your original performance sheet. A degradation of more than 10–15% from design strongly suggests fouling. A 1 mm ash deposit on convective tubes can cut the local heat transfer coefficient by 15–30% and push stack temperature up by 20–50°C — numbers that compound fast on a unit running 8,000 hours a year.
Physically, open access doors on the convective pass and inspect. Light, powdery gray ash is usually easy to dislodge with soot blowing. Hard, sintered or sulfated deposits — darker, sometimes glassy — require mechanical rapping or high-pressure water washing offline. Note the deposit color and hardness in your maintenance log; this pattern over multiple inspections tells you whether your soot blowing frequency is adequate.
Acoustic soot blowers, where installed, can run online during operation. Mechanical soot blowing typically requires a short load reduction. If you’re not already tracking soot blowing cycles versus stack temperature response, start now — it’s one of the fastest ways to optimize cleaning intervals without unnecessary shutdowns.
Waterside Scaling Diagnosis
Pull a boiler water sample and test TDS, conductivity, and hardness at least weekly during normal operation — more often if your makeup water quality varies seasonally, which it does in many plants drawing from surface water sources. A calcium carbonate scale layer of just 0.5 mm on the waterside tube surface is estimated to increase equivalent heat loss by roughly 2–3%, because calcium carbonate’s thermal conductivity is orders of magnitude lower than steel.
If conductivity is trending upward over weeks, your blowdown regime is insufficient. If you’re seeing elevated hardness, your softener or RO system needs attention before the boiler does.
When to Clean Online vs. Offline
| Condition | Recommended Action | Typical Downtime |
|---|---|---|
| Stack temp elevated 10–20°C, powdery ash deposits | Increase soot blowing frequency; online acoustic blowing | Zero |
| Stack temp elevated 20–40°C, moderate sintered deposits | Scheduled offline mechanical soot blowing + gas-side water wash | 8–16 hours |
| Stack temp elevated >40°C, hard sulfated deposits | Planned outage, chemical cleaning or manual descaling | 24–48 hours |
| Waterside TDS >3× design limit, visible scale in sample | Offline chemical descaling (inhibited acid wash) | 24–72 hours depending on circuit volume |
Turnaround hours vary considerably depending on boiler size, deposit severity, and whether your team is doing the work in-house or calling in a service crew. Don’t let an overly optimistic schedule push you to shortcut the post-cleaning flushing and water quality verification — coming back online with residual acid or loose scale debris causes more problems than the fouling did.

Identifying and Repairing Tube Corrosion, Erosion, and Leakage in WHRB Heat Surfaces
Tube failure is, in my experience, the failure mode that actually takes a WHRB offline — not controls faults, not feedwater issues, but metal loss you didn’t catch until a pinhole became a rupture. The three mechanisms behave very differently, so treating them as a single “corrosion problem” leads to wrong repairs and fast recurrence.
High-Temperature Sulfur Corrosion and Acid Dew Point Attack
High-temperature sulfur corrosion typically targets superheater and high-temperature evaporator tubes when flue gas carries SO₃ and inlet temperatures run above roughly 500°C. The attack is chemical — sulfate compounds react with oxide scale, producing a molten or semi-molten layer that dissolves the tube wall progressively. You’ll see it most in WHRBs sitting behind sulfur-bearing fuel combustion: sulfuric acid plants, non-ferrous smelting off-gas, certain process furnaces. Wall loss is uneven, often worst on the windward side of the tube or at weld toes where scale discontinuities exist.
Low-temperature acid condensation is the opposite end of the same problem. On economizer tubes — especially the cold-end bundles — whenever the tube wall temperature drops below the acid dew point (typically 120–150°C for flue gas with meaningful SO₃ content; the exact threshold depends on SO₃ concentration and moisture partial pressure), sulfuric acid condenses directly on the metal surface. The result is a pitting pattern on the outer tube wall, often masked by ash deposit until an inspection reveals orange-brown staining or pitting craters beneath.
For flue gases with SO₃ concentrations above roughly 20 ppm, the acid dew point can reach 140–155°C, which means economizer outlet gas temperatures below that range will cause active condensation corrosion on standard carbon steel tubes.True
SO₃ reacts with moisture to form H₂SO₄ vapor; the dew point rises sharply with SO₃ concentration. Standard references (VGB, ASME) confirm this threshold band, and it is why ND steel (09CrCuSb) was specifically developed for cold-end economizer applications.
Oxygen pitting is a waterside problem that gets overlooked during scheduled outages. When a WHRB sits idle with the drum partially filled and no nitrogen blanketing or chemical treatment maintaining oxygen scavenger residual, dissolved oxygen attacks the tube interior — especially in lower headers and economizer coils where water stagnates. The pits are hemispherical, localized, and can perforate a tube wall in a fraction of the time external corrosion would.
Erosion in Dust-Heavy Applications
Cement kiln waste gas, electric arc furnace off-gas, and certain steel plant WHRBs carry particulate loadings that make erosion the dominant wall-loss mechanism rather than corrosion. Gas velocity matters enormously here. At convective tube banks, local velocity above roughly 12–15 m/s accelerates impingement erosion sharply — tube wall thinning in the range of 0.1–0.5 mm per year in severe cases, with the actual rate depending on particle hardness, concentration, and impingement angle. Leading-edge tubes in the first row of each bank suffer worst. Wear is not uniform; you’ll often find one tube in a row reduced to minimum wall while its neighbor is still near nominal thickness.
Field Detection Methods That Actually Work
Ultrasonic thickness testing (UT) is the workhorse. During planned outages, map UT readings across erosion-prone locations — first and last rows of tube banks, bend zones, areas directly downstream of flow disturbances like baffles or sootblower paths. A systematic grid with readings logged against a baseline from commissioning will reveal trend lines before you hit minimum wall. Don’t rely on visual inspection alone; a tube with 0.8 mm remaining can look perfectly intact under a layer of ash.
Dye-penetrant (PT) inspection targets weld toes at tube-to-header joints, where stress concentration and corrosion interact. Cracks initiate here before through-wall leakage develops.
Online acoustic emission (AE) monitoring, where installed, picks up active micro-leaks as stress waves propagate through the pressure part — useful for detecting slow seepage between outages without requiring shutdown.
Tube leak symptoms in operation are usually not subtle once you know what to look for: wet ash patches or steam wisps at the casing inspection doors, unexplained drum level drop that the feedwater control keeps chasing, a step-change in flue gas humidity on the stack analyzer, or a steam flow imbalance between circuits that wasn’t there last week. Any one of these warrants an immediate load reduction and inspection.
Repair Decision Logic
The repair decision depends on how widespread the damage is, not just on a single bad tube.
| Damage Extent | Recommended Action | Caveat |
|---|---|---|
| Isolated pinholes, ≤2–3 tubes, wall loss localized | Plug affected tubes, schedule segment replacement at next outage | Monitor adjacent tubes closely; isolated failures in corrosive service rarely stay isolated |
| Clustered failures in one zone, 5–15% of bank | Segment replacement during planned shutdown | Investigate root cause first — replacing tubes into the same corrosive or erosive condition repeats the cycle |
| >15% of tubes below minimum wall thickness | Full bank replacement | Partial repair at this density is usually false economy |
Material selection on replacement matters more than most procurement teams realize. For cold-end economizer tubes in SO₃-bearing service, ND steel (09CrCuSb) is the established choice — its copper and chromium additions raise the corrosion resistance in dilute sulfuric acid significantly compared to plain carbon steel. For superheater sections operating above roughly 550°C metal temperature, T91 or P91 (9Cr-1Mo-V) alloy is appropriate. Where erosion is the primary threat — leading-edge tubes in cement or steel WHRB applications — ceramic coating sleeves or sacrificial wear shields on the upstream face of the first tube row can extend service intervals by 2–4×, depending on the particle load. It’s an unglamorous fix, but it works.
Troubleshooting Steam Drum Level Instability, Carryover, and Circulation Faults
Steam drum problems are where a lot of WHRB headaches actually live — and they’re frequently misdiagnosed because the root cause is upstream of the drum itself. The drum is just where the symptom shows up.
Natural Circulation in WHRBs: Why Low-Temperature Units Are Especially Vulnerable
Natural circulation in any boiler relies on one thing: the density difference between the water in the downcomers (cooler, denser) and the steam-water mixture rising through the risers (lighter). That density gradient is your driving head. In a conventional fired boiler with furnace heat fluxes well above 100 kW/m², that head is substantial. In a WHRB recovering flue gas at 350–400°C, you’re working with much lower heat input per unit area, and circulation ratios can drop to 5:1 or even tighter in some low-pressure designs — compared to 15:1 or higher in high-pressure utility boilers.
That low ratio means the system is genuinely sensitive to flow imbalances. If one bank of riser tubes gets slightly more fouled than another, or if a stagnant zone develops during low-load operation, you’re close to the edge. Partial dry-out in riser tubes isn’t a theoretical risk — it happens, usually silently until tube metal temperatures start climbing or you see a tube failure during the next inspection.
Drum Level Instability: Four Causes, One Misleading Display
The most common scenario: operators see the drum level swinging ±50–80 mm on the display, assume the feedwater control valve is hunting, swap the valve internals, and the problem continues. Nine times out of ten in a cement or metallurgical plant, the real culprit is the level measurement itself.
Dusty, high-particulate environments foul the impulse lines on differential pressure transmitters — slowly, over weeks, until the reading lags reality by several minutes or drifts permanently. The fix is routine impulse line flushing (quarterly at minimum in cement applications) and cross-checking three independent readings: the magnetic float gauge, the DP transmitter output, and the direct gauge glass. If those three disagree by more than 20–30 mm under steady-state conditions, trust the gauge glass and start tracing the instrumentation fault before touching anything else on the control side.
The swell-and-shrink effect adds another layer of confusion. During a sudden load increase, drum pressure momentarily drops, the water flashes partially, and the level indicator reads high — even though actual water inventory is unchanged or falling. Operators responding to that false high level by cutting feedwater can spiral into a genuine low-water event within minutes. Any control loop tuned without accounting for this dynamic will hunt.
Differential pressure transmitter impulse line blockage is a frequent cause of drum level instability in dusty industrial WHRB installationsTrue
Particulate-laden flue gas environments such as cement kilns and electric arc furnaces allow fine dust to migrate into instrument impulse lines over time, causing DP readings to drift or freeze. This is well-documented in field maintenance practice and is distinct from the drum level dynamics themselves.
Steam Carryover: What the Turbine Tells You First
Carryover is one of those faults the turbine reports before the boiler instrumentation does. You’ll see it as unexplained erosion on first-stage blading, or steam purity trending upward — SiO₂ or Na⁺ climbing, cation conductivity pushing above the 0.2 µS/cm threshold that most operators use as their alarm point. Superheater outlet temperature suppression is another early signal: wet steam absorbs superheat rapidly, so the outlet temperature drops 10–20°C below setpoint without any obvious reason.
When carryover is confirmed, reduce boiler load by 15–20% first. That alone often improves separation enough to stabilize purity while you investigate further. Then open the drum for internal inspection at the next available outage. Chevron separators and cyclone demisters crack, warp from thermal cycling, or simply get pushed out of position by years of vibration. A bent chevron vane won’t show on any online instrument — you have to look at it.
Blowdown discipline matters here too. Continuous blowdown rate should be calibrated against actual water chemistry, not set-and-forgotten. Surface blowdown removes the dissolved solids concentrated at the waterline; skipping it for weeks builds up the salinity gradient that drives carryover regardless of separator condition.
Forced Circulation Units: Pump and Flow Meter Faults
Pump-assisted WHRBs — common in horizontal configurations or where heat flux is too low to sustain reliable natural circulation — introduce their own failure modes. Cavitation is the one that causes the most confusion because it sounds like mechanical vibration, which operators sometimes attribute to pipe supports or thermal expansion. If the suction pressure at the circulation pump drops below the fluid’s saturation pressure (easy to happen if suction strainers partially block), you’ll get intermittent cavitation that shows as erratic flow meter readings and noise before it progresses to impeller damage.
Check the bypass valve on the pump discharge circuit for internal leakage. A passing bypass valve is common after two or three years of operation and reduces effective circulation flow by 10–25% without triggering any obvious alarm — the flow meter reads total pump output, not net flow to the risers. Verifying actual tube-side flow requires either ultrasonic clamp-on measurement during shutdown preparation or a careful pressure drop comparison against the original commissioning baseline.
In practice, forced-circulation WHRB reliability lives or dies on pump maintenance intervals and strainer inspection frequency. Six-month strainer checks are the minimum; quarterly is better in dirty flue gas service.
Resolving Slagging, Ash Bridging, and Gas Flow Maldistribution in the Flue Gas Path
Ash deposition problems are, in my experience, where most WHRB operators lose the most ground — quietly, over weeks, until stack temperatures have climbed 40°C and nobody can explain exactly when it started. Unlike a tube leak, which announces itself, slagging and flow maldistribution erode performance gradually and only become obvious during an outage inspection when you’re staring at a tube bank that looks like a cave formation.
Ash Fusion Temperature: Why Cement Kiln Slag and Steel Mill Scale Are Not the Same Problem
The first thing to establish is what kind of ash you’re dealing with, because the remediation strategy differs significantly. Clinker-carry dust from cement kilns has an ash fusion temperature of roughly 1,200–1,350°C. By the time that material contacts your first convective tube rows — where gas temperatures might still be 850–950°C — it’s arriving partially molten and sticky. It bonds to tube surfaces aggressively and, once a thin layer forms, insulates the tube enough to raise the local surface temperature further, encouraging more deposition. Steel mill scale is a different character: fusion starts lower, around 800–1,100°C depending on the specific furnace chemistry, but it tends to form harder, more brittle deposits that are actually easier to mechanically dislodge once you get to them.
Alkali content is the variable that catches plants off guard. When K₂O and Na₂O combined exceed roughly 3% by mass in the ash, the effective stickiness threshold drops substantially — sometimes by 100–150°C — meaning deposits form at gas temperatures you’d normally consider safe. If your fuel or raw material supply has shifted (new limestone source, different coal blend in the steel mill), get a fresh ash composition analysis before assuming your existing soot blower schedule is still adequate.
Gas Flow Maldistribution and Why It Compounds Everything
Partial ash bridging between tube rows — even a 20–30% lane blockage — doesn’t reduce heat transfer uniformly. It forces the remaining gas volume through narrower channels, increasing local velocity and particle impact energy. Localized erosion rates in these channels can run 3–5× higher than the design baseline, meaning tube walls in the channeled zone thin out years ahead of schedule while adjacent tubes remain essentially untouched. Eroded or misaligned baffle plates accelerate this further by directing high-velocity ash-laden gas at the same tube sections repeatedly.
Expansion joint condition matters more than most maintenance schedules acknowledge. An expansion joint that has partially collapsed or shifted skews the duct cross-section, creating an asymmetric velocity profile before the gas even enters the tube bank. Check joint gap uniformity at both planned outages and, where accessible, during operation using thermal imaging — a temperature differential of more than 30–40°C across the same tube row is a reliable sign of maldistribution.
Soot Blower Audit: The 40% Problem
Inadequate soot blowing accounts for approximately 40% of WHRB efficiency losses in cement kiln waste heat applications.True
This estimate is consistent with field performance data from cement WHRB operators and is referenced in industry maintenance benchmarking for high-dust WHRB installations where continuous ash loading is high and blowing intervals are not dynamically adjusted.
Long retractable soot blowers require steam at 1.0–1.6 MPa to generate sufficient jet momentum — lower than that and you’re just warming the deposit without dislodging it. Check nozzle wear at every outage; worn nozzles increase jet divergence and cut effective blowing radius by 30–40%. Travel speed matters too: blowers moving too fast don’t dwell long enough on stubborn deposits; too slow and they cause thermal shock cracking on already-stressed tubes. Blowing sequence programming should work upstream-to-downstream so dislodged ash from the first rows doesn’t redeposit on already-cleaned downstream surfaces.
Cold-End Hopper Plugging
At the cold end of the gas path, ash cools, gains cohesion, and bridges. Hopper half-angles below 60° from horizontal are a chronic offender — ash arches across the outlet and backfills into the gas path, partially blocking the last tube rows. Rotary valve failures or drag chain conveyor jams make this worse quickly. In colder climates or seasonal shutdowns, even a brief drop in hopper wall temperature can cause a deposit that takes hours to clear manually.
During planned outages, the inspection protocol should include rope access or scaffolding survey of actual tube lane spacing (compare against design drawings — bridging often reduces clear width by 15–40 mm), baffle plate alignment verification with a straightedge or laser level, and direct measurement of expansion joint gaps at multiple points across the duct width.
Design-Level Fixes for Chronic Slaggers
If you’re seeing repeat slagging despite optimized soot blowing, the boiler design itself may be mismatched to the ash load. Wider longitudinal tube pitch — above 100 mm is a useful starting threshold — reduces the bridging probability significantly. Aerodynamic tube shields on the leading-edge tubes of the first rows deflect particle impact away from the weld-affected zone. Online vibration rapping systems, common in European cement WHRB installations, apply periodic mechanical impulse to tube panels and can maintain heat transfer coefficients 10–20% higher than equivalent soot-blowing-only installations in high-alkali applications. These are not cheap retrofits, but against the cost of a tube bank replacement every three to four years, the economics usually work out.

Diagnosing Feedwater, Blowdown, and Water Treatment Problems That Damage WHRB Internals
Water chemistry is where a lot of WHRB damage originates — quietly, invisibly, over months — and it’s consistently underestimated by plant teams focused on the mechanical side. By the time you see tube pitting or scale-induced hot spots, the chemistry problem has usually been running for weeks. Catching it earlier requires systematic monitoring, not just periodic lab samples.
Water Quality Targets That Actually Matter in Service
The numbers matter and so do the dependencies behind them. Feedwater hardness should stay below 0.03 mmol/L — that’s non-negotiable regardless of pressure rating, but the consequences of exceeding it scale sharply with operating pressure. At low-pressure units (under 1.6 MPa), soft scaling builds slowly. Above 3.82 MPa, even brief hardness excursions can cause rapid carbonate and silica deposition on superheater tubes where temperatures are highest and flow velocities lowest.
Dissolved oxygen (DO) at deaerator outlet should be below 7 µg/L under normal operation. Boiler water pH should be held between 9.0 and 10.5 — too low and you get acidic corrosion; too high, particularly above 11, and you risk caustic gouging in high heat-flux zones. TDS limits are pressure-dependent: below 2000 mg/L for units operating under roughly 1.0 MPa, tightening to below 500 mg/L for anything above 3.82 MPa. Pushing beyond these limits while running continuous blowdown “as usual” is not a safe workaround — it’s a sign the makeup water treatment upstream isn’t doing its job.
Deaerator Performance: The First Place to Check
A deaerator running poorly is often the hidden root cause of economizer tube pitting. If DO at the deaerator outlet creeps above 15 µg/L, that’s a red flag. Common causes are a sticking or undersized vent valve (the vent needs to pass a small but continuous steam flow to sweep out non-condensable gases), inadequate heating steam supply, or — and this one gets missed frequently — feedwater arriving at the inlet well below saturation temperature for the operating pressure. The inlet water must reach saturation temperature inside the deaerator to release dissolved gases efficiently. If your makeup water is cold or the heater upstream is underperforming seasonally, you won’t hit that condition and DO will creep up.
Check DO with an inline sensor, not just manual titration. Titration gives you a snapshot; inline sensors show you the drift over a shift.
Feedwater Pump Cavitation
Cavitation in feedwater pumps is a mechanical problem with a chemistry link. High feedwater temperature — particularly when deaerator backpressure is low or feedwater has absorbed heat from recirculation — reduces the NPSH margin and causes cavitation. You’ll hear a characteristic rattling or crackling sound, see erratic discharge pressure, and eventually find impeller surface pitting on inspection. Low-pressure WHRBs with compact piping layouts are especially prone to this because the suction line geometry is sometimes compromised to fit the layout. The fix usually involves checking NPSH available versus pump curve at actual operating temperature, not design temperature — those two numbers diverge in summer months when cooling water temperatures rise.
Blowdown Strategy: The Energy-Cost Trade-off
Continuous blowdown rate should be calculated from feedwater TDS and the maximum allowable boiler water TDS, not set arbitrarily.True
The correct blowdown ratio equals feedwater TDS divided by (allowable boiler water TDS minus feedwater TDS). Setting blowdown by feel or habit leads to either energy waste or accelerated scaling.
Over-blowing wastes roughly 1–3% of total steam energy — that’s real fuel cost, and in a WHRB where you’re already recovering waste heat, throwing it away through excessive blowdown is doubly wasteful. Under-blowing lets TDS climb, scale forms on tube surfaces, and localized boiling under the scale layer creates the exact conditions for stress corrosion. Run the calculation properly. Use a blowdown heat recovery vessel if continuous blowdown rate exceeds roughly 2–3% of steam output; payback is usually under two years at typical fuel or waste heat valuations.
Chemical Dosing Faults
Phosphate dosing errors are particularly common in units that run intermittently or at variable load. At high temperatures and pressures, phosphate can precipitate out of solution — “phosphate hideout” — which temporarily depletes the buffer while the boiler is hot, then releases back into solution when load drops, causing a pH spike. If your boiler water pH swings are happening at load changes rather than steadily, phosphate hideout is the likely mechanism.
Oxygen scavenger underdosing — sodium sulfite for lower-pressure units, hydrazine or alternatives for higher-pressure steam — leaves residual DO that attacks economizer tube surfaces from the waterside. The resulting pitting is small but deep, and it’s often found during outage inspections when the tube has already thinned locally.
Silica deserves specific attention in plants drawing makeup water from groundwater sources or surface water in certain regions. Silica carryover produces white, glassy deposits on superheater tubes. It’s hard to remove mechanically and increases tube wall temperature by insulating the surface. If your makeup water silica exceeds roughly 20 mg/L, you need either reverse osmosis pretreatment or strict pH and blowdown management to keep boiler water silica below the carryover threshold for your operating pressure.
Connecting the chemistry symptoms to observable evidence: pitting concentrated in the economizer inlet section points to DO control failure; white chalky or glassy deposits on superheater tubes point to silica carryover; black magnetite sludge accumulating in the drum bottom or visible in blowdown is usually a sign of pH excursion — either a temporary acidic upset or a period of caustic concentration. Each of these symptoms has a different corrective path, and treating them all as “a water quality problem” without distinguishing the mechanism wastes time and doesn’t protect the heat surfaces.
Troubleshooting Instrumentation, Control Loops, and Safety Valve Malfunctions in WHRB Systems
Instrumentation failures are, in my experience, the single most misdiagnosed category in WHRB troubleshooting. A plant team pulls maintenance on heat transfer surfaces, replaces tubes, adjusts sootblowers — and the problem persists. Three weeks later someone checks the drum level transmitter impulse line and finds it half-blocked with condensate and rust scale. The boiler was never actually misfiring. The control system was lying.
The Instrumentation Loop Architecture You Need to Know Cold
A properly instrumented WHRB carries a predictable set of critical loops: drum pressure transmitter, three-element drum level control (drum level + steam flow + feedwater flow), flue gas inlet and outlet temperature sensors, steam flow metering, and superheater outlet temperature — on superheated steam configurations operating above roughly 350–450°C. Each of these loops has its own failure signature, and conflating them wastes days.
The drum pressure transmitter is your reference for nearly every protection function. If it drifts, your safety valve actuation setpoint, your pressure-based level correction, and your steam flow inference all drift with it. Impulse lines on pressure transmitters are particularly vulnerable in plants located in cold climates — ambient temperatures below -5°C can partially freeze condensate pots, introducing a slow, creeping offset that shows up as a gradually rising apparent pressure reading with no corresponding process change. The fix is heat tracing on impulse lines, which is standard in northern China, Russia, and Central Asian installations but still missing in some tropical plants that get relocated or duplicated to colder sites without re-engineering the field instrumentation package.
Thermocouple drift in high-dust flue gas ducts deserves its own conversation. Type K thermocouples above 900°C will typically show 10–30°C of positive drift after roughly 4–6 months of continuous service, driven by grain boundary migration in the Chromel element at sustained high temperatures. This matters because your flue gas inlet temperature reading feeds both the efficiency calculation and, on many installations, the high-temperature trip that protects the first superheater pass. A sensor reading 20°C low means your protection margin is eroding silently. Replace or cross-check against a reference portable thermocouple annually — or more often if the upstream process (cement kiln, electric arc furnace) runs hotter than the original design case.
Vortex flow meters on steam lines are reliable under stable conditions but susceptible to zero drift after maintenance, particularly if the meter body is reinstalled with any flow disturbance within 10–15 pipe diameters upstream. A zero drift of 2–5% on steam flow reads directly into your three-element controller’s feedforward signal, which then biases feedwater valve position continuously.
Three-Element Controller Tuning: Where Most Startups Go Wrong
The three-element drum level controller is elegant in steady state and fragile during transients. During cold startup, when steam flow is low and drum level is changing rapidly due to thermal expansion, the integral term accumulates error faster than the process can respond — classic integral windup. The result is feedwater valve hunting: the valve slams open, overshoots drum level, then closes hard, creating the swell-and-shrink oscillation that operators sometimes mistake for a circulation fault.
Correct practice is to start in single-element mode (drum level only) during low-load startup, then switch to three-element once steam flow exceeds roughly 20–25% of rated capacity and the signal-to-noise ratio on the flow meters is acceptable. The switchover logic needs to be explicitly programmed and tested, not assumed. Many DCS configurations I’ve seen in field commissioning leave the switchover threshold at factory default, which is often wrong for the specific boiler’s drum volume and ramp rate.
PID tuning for the feedwater valve should be done with the actual installed valve, not on the design specification. Control valve hysteresis in a worn globe valve can be 3–8%, which means aggressive integral gain just produces limit cycling. If you’re seeing persistent ±50 mm level oscillation at steady load, check valve hysteresis before touching PID parameters.
Three-element drum level control is required by most industrial boiler standards for WHRBs above a certain capacity threshold.True
ASME, EN 12953, and most national boiler inspection codes require three-element feedwater control for boilers above a specified steam output (commonly around 10 t/h and above), precisely because single-element control cannot adequately compensate for the shrink-and-swell phenomenon during load changes.
Safety Valve Chattering and Weeping: Not Just an Annoyance
A safety valve that weeps — leaks a thin film of steam at normal operating pressure — is usually telling you one of two things: the set pressure is too close to maximum allowable working pressure, or the seat has been wire-drawn by previous steam cuts. ASME Section I and EN 13445 both require a minimum 3% margin between operating pressure and safety valve set pressure. In practice, many WHRB installations creep up their operating pressure over time (chasing a little more output) without rechecking safety valve set points. The valve starts to simmer. Each simmer cuts the seat slightly. Within a few months you have a valve that weeps at 80% of set pressure.
Chattering — rapid, repetitive opening and closing — is a different problem, usually caused by oversized valve capacity relative to the boiler’s actual relief demand, or excessive back pressure on the discharge stack. Back pressure above roughly 10% of set pressure on a conventional (non-balanced bellows) safety valve will cause instability. If your discharge stack was extended after original installation, this is worth checking.
Test your safety valves by hand-lift annually at minimum, and by actual lift-to-set-pressure test every 2–3 years depending on regulatory jurisdiction. Document response time.
DCS and Trip System Verification
Distinguishing a genuine process excursion from an instrumentation fault requires historian data. A real high-pressure event will show correlated drum level response, steam flow change, and feedwater valve movement. An instrumentation fault typically shows an isolated signal spike with no correlated response in adjacent measurements. Two-out-of-three voting logic on critical trips — high drum pressure, low drum level, high flue gas temperature — exists precisely to filter spurious sensor failures from genuine trips, but only if all three sensors are independently calibrated and their impulse lines are independently routed. I’ve seen plants where two of the three level transmitters shared a common impulse line manifold. That is not two-out-of-three protection; that’s one-out-of-one with extra wiring.
Annual functional testing of every interlock and trip, with documented response times compared against design specification, is the baseline. If your SIL assessment assigned SIL 2 to the low-level trip, your proof-test interval is probably 1–2 years — honor it, and keep the records.

Refractory, Casing, and Structural Integrity Issues: Detection and Remediation
Structural and refractory failures rarely announce themselves dramatically. They creep — a slightly warm panel here, a hairline weld crack there — until one day you’re looking at a warped casing section, a flooded expansion joint cavity, or a waterwall tube that’s baked itself into failure because the hanger system stopped allowing free movement three shutdowns ago. In WHRBs, this category of problem gets underattended precisely because it doesn’t trip an alarm.
The Casing Hot-Spot Phenomenon
When internal refractory or insulation lining spalls — and in high-alkali environments like cement kiln bypass gas streams, it eventually will — the outer casing steel loses its thermal buffer. Surface temperatures that should sit at 40–60°C on a properly insulated panel can climb to 200–400°C, depending on flue gas temperature and how much lining has dropped away. At those temperatures, the casing steel begins to distort. Welds that were designed for ambient-temperature loads start seeing cyclic thermal stress they were never sized for. External oxidation accelerates sharply above roughly 250°C on carbon steel, and what started as a lining problem becomes a structural repair job.
The insidious part is that a 0.5 m² spall patch on the inside of a large WHRB can read as a hot spot no bigger than a dinner plate on the outside — easy to dismiss on a casual walkdown.
Conducting an Infrared Thermography Survey
IR thermography during operation is the right tool here, but it needs to be done systematically, not just with a point-and-shoot pass on the way to lunch. Walk every accessible casing panel in a grid pattern — overlapping frames, consistent camera angle, ideally with the unit at stable load for at least two hours before scanning. Wind on the casing exterior will cool the surface and mask real hot spots; do this on calm days or shield the panels during scanning.
The benchmark that most experienced inspection teams use: a surface temperature above 80°C on an insulated WHRB panel is a reliable indicator that the lining behind it has degraded by 50% or more in that zone. Above 120°C, treat it as full lining loss until proven otherwise. Log GPS or grid coordinates of every anomaly — you need a map, not just photos, so the repair crew can find the zones during the offline window.
An insulated WHRB casing panel reading above 80°C on IR thermography typically indicates at least 50% lining failure behind that panel.True
This threshold is consistent with heat transfer calculations for standard ceramic fiber or castable refractory lining thicknesses used in WHRB construction; partial lining loss reduces thermal resistance proportionally, and surface temperature rise becomes measurable at that degradation level.
Expansion Joint Failures and False Air Infiltration
WHRBs connected to cement kiln preheater towers or glass furnace regenerators deal with 10–30 mm of differential thermal expansion at connection points — the actual range depends on duct length, material, and operating temperature delta. Fabric expansion joints handle that movement, but fabric deteriorates faster in alkali-laden or high-particulate gas streams. Metallic bellows corrode from the outside if insulation jacketing is breached.
A failed expansion joint has two consequences that directly hurt performance. First, false air infiltrates the flue gas path, diluting the hot gas stream and dropping the temperature available to the WHRB heat surfaces — this can trim 15–30°C off inlet temperature, which matters when you’re already operating at the low end of the efficiency window. Second, dust leaks outward, creating housekeeping and sometimes safety issues at grade. Check expansion joints visually and with a smoke pencil at the joint perimeter during operation. Any visible dust pluming or inward flutter on the fabric is a failure.
Refractory Spalling: Why It Happens in WHRB Applications
Thermal shock is the primary driver — rapid startup and shutdown cycles that WHRB operators sometimes treat as routine. Alkali attack from K₂O and chloride-rich gases in cement kiln bypass streams chemically degrades high-alumina castable over time, causing subsurface cracking that isn’t visible until a chunk drops. Induced draft fan pulsation, particularly on units with variable-frequency drives hunting at part load, introduces vibration that works existing microcracks open. A combination of two or three of these factors accelerates spalling dramatically faster than any single cause would suggest.
Repair Planning: Emergency Patching Versus Full Reline
For spall zones under roughly 0.3–0.5 m², castable refractory patch repair during a short planned outage is usually viable — anchor the surrounding lining, clean the substrate, apply a compatible castable, and cure it properly before restart. Skipping the cure cycle is the most common mistake; uncured castable exposed to rapid heat-up will spall again within weeks.
For failures exceeding about 2 m², a full panel reline is the only durable answer. Budget 5–10 days of offline time depending on WHRB size, access scaffolding complexity, and whether the refractory anchor system needs replacement. In practice, plants often discover during a partial reline that adjacent zones are also compromised — scope creep on refractory jobs is the rule, not the exception.
Structural Framework: Buckstays, Hangers, and Nozzle Welds
This gets overlooked even by experienced maintenance teams. Check buckstay alignment during every planned outage — misaligned buckstays bind the waterwall panels and prevent the thermal expansion the design intended. Hanger rods should be loaded evenly; a rod that’s gone slack means load has transferred somewhere else. Spring hangers that have bottomed out or coil-bound are a clear sign the system has moved beyond its design range and isn’t moving back.
Improper thermal expansion restraint is a root cause of waterwall panel cracking and nozzle weld failures, both of which eventually mean tube leaks. It’s a slow mechanism, but the accumulated damage from three or four years of constrained expansion is very real. Verify spring hanger travel and load indicator positions against the original design data during every major inspection.
Preventive Maintenance Schedule and Condition Monitoring Program for Long-Term WHRB Reliability
All the diagnostic work covered in earlier sections is reactive by nature. You find the problem after it has already cost you something — downtime, scrap steam, a tube replacement, or worse. A structured PM program shifts that balance. It will not eliminate every failure, but it compresses the gap between “something is degrading” and “we caught it before it became a shutdown.”
Daily Operator Rounds: The Foundation Nobody Respects Enough
Operators walking the floor twice a shift, logging stack temperature, drum pressure, feedwater flow, and blowdown conductivity takes maybe 20 minutes. In practice, many plants treat this as a checkbox exercise. That is a mistake. Stack temperature trending is your primary efficiency KPI — a steady 15–20°C rise over two or three weeks almost always signals tube fouling building ahead of a measurable steam output drop. Operators who are trained to trend readings rather than just record them will catch that signal before it degrades your heat transfer coefficient enough to matter.
A simple log sheet — time-stamped, with a flagging column for out-of-range values — is sufficient. No sophisticated software required at this tier.
Monthly Tasks: Where Condition Monitoring Earns Its Keep
Monthly inspections should cover soot blower operation (confirm each lance travels its full stroke, check steam pressure differential across nozzles), vibration spot-checks on ID fans and feedwater pumps with a handheld analyzer, and a water chemistry review against your treatment program baseline. Conductivity, pH, silica, and dissolved oxygen should all be trended against control limits, not just compared to a single target number.
This is also the right interval to review your DCS historian for any control loop that has been running in manual or that shows hunting behavior — a drum level controller oscillating with ±80 mm swings, for instance, usually points to a tuning issue or a transmitter that has drifted.
Annual Planned Outage: Non-Negotiable Scope
Annual outages are where you protect the long-term asset. Minimum scope for a cement kiln or steel furnace WHRB should include tube UT survey on high-velocity gas path zones (particularly the convective bank inlet rows), safety valve bench testing and re-certification, full refractory visual inspection with hammer-tap survey on cast sections, and a heat transfer performance test benchmarked against the original design pinch point — ideally within 5°C of design value. A pinch point deviation of 10°C or more usually means fouling or reduced gas flow distribution that the daily logs missed.
Before returning to service after any chemical cleaning (a NaOH and Na₃PO₄ boil-out at roughly 0.3–0.5% concentration is standard for new units or after significant scale removal), run a hydrostatic test at 1.25× design pressure, hold for the code-required duration, and confirm no weeping at rolled tube ends or at welded nozzle connections.
Spare Parts Criticality: Two Tiers, One Decision
| Class | Items | Strategy |
|---|---|---|
| A — Hold on-site | Safety valves, drum level transmitters, feedwater control valve trim, circulation pump mechanical seals | Zero lead-time tolerance; replace at PM, return failed unit for repair |
| B — 2–4 week lead time acceptable | Tube section blanks (pre-bent to drawing), soot blower lances, thermocouple assemblies | Confirm supplier stock before each outage window |
World-class WHRB availability in continuous cement plant service reaches 92–96% when structured PM and condition monitoring are maintained.True
This range is consistent with published cement industry performance benchmarks and is achievable with rigorous PM, proper water chemistry control, and annual planned outage scope. Plants without structured programs typically operate in the 82–88% range.
Condition Monitoring Technology Integration
Continuous stack temperature monitoring piped into your DCS historian gives you a rolling efficiency picture without any manual intervention. Add thermal imaging on casing panels quarterly — hot spots above roughly 80°C surface temperature indicate refractory void or bypass gas channeling, both of which cause accelerated external tube oxidation. Online steam purity analyzers (cation conductivity, sodium) on the steam header are worth the investment if you are supplying steam to a process sensitive to contamination.
Acoustic leak detection during normal operation can localize tube failures to within a few meters on a long convective bank, which compresses outage repair time significantly compared to finding it by brute-force visual inspection during a cold boiler walk.
Our after-sales team provides remote DCS data review and annual performance audit services for international EPC projects, with spare parts supply chains structured specifically for sites where in-country sourcing is limited. For plants running in remote locations — West Africa, Central Asia, parts of Southeast Asia — that supply chain certainty is often more operationally important than the equipment price itself.

Frequently Asked Questions About Waste Heat Recovery Boiler Troubleshooting
How do I know if my WHRB efficiency has degraded before an alarm trips?
Three trends usually appear well before any DCS alarm activates. First, watch your stack temperature — a steady climb of 15–30°C over four to eight weeks at constant upstream process load almost always points to fouling on convective surfaces. A 1 mm ash deposit layer alone can push stack temperature up by 20–50°C and strip 15–30% off the effective heat transfer coefficient. Second, if steam output is falling while the kiln or furnace throughput hasn’t changed, that mismatch is the clearest early signal. Third, calculate your pinch point temperature difference regularly; when it starts drifting upward from its design value — even 8–12°C of creep — fouling or circulation problems are the likely culprits. Most operators only look at alarm states. The real diagnostic work happens in the trends.
What is the typical service life of WHRB heat exchange tubes, and when should I plan replacement?
Carbon steel tubes in reasonably clean flue gas — say, a glass furnace or a coke oven gas application — commonly reach 15–25 years before wall thinning becomes a structural concern. Drop that to 8–12 years in cement kiln or high-sulfur chemical reactor exhaust, where alkali attack and dew-point corrosion accelerate external degradation significantly. The trigger for planning a full tube replacement isn’t a single failure; it’s a systematic UT thickness survey showing more than roughly 20% of tubes in any bundle section below minimum design wall thickness. Spot repairs on a heavily thinned bundle buy time, not reliability. Budget your replacement campaign before you hit that threshold, not after the first blowout.
Can a WHRB be retrofitted if upstream process capacity expands?
Yes — but it’s not a damper swap and done. Higher flue gas flow means higher gas-side velocity, which immediately raises erosion risk on existing tube surfaces. Depending on how large the capacity jump is (anything above roughly 15–20% increase warrants a full thermal-hydraulic recalculation), you may need to add heat surface modules, upsize bypass damper arrays, or upgrade the steam drum if the new steam generation rate exceeds original drum disengagement capacity. Every one of those changes touches pressure vessel certification — ASME Section I, PED, or GB 150, depending on jurisdiction. Factor in lead time for third-party inspection re-certification. In practice, a retrofit that looks straightforward on paper often surfaces drum nozzle sizing limitations that nobody anticipated.
What is the difference between a WHRB and an HRSG, and does troubleshooting differ?
The terminology gets blurred in procurement documents constantly. An HRSG is specifically designed for gas turbine exhaust: relatively clean, high-temperature gas, predictable flow, low particulate loading. WHRBs handling cement kiln, EAF, or nonferrous smelter exhaust deal with dust concentrations that can run 20–80 g/Nm³ or higher, variable gas temperatures, and chemically aggressive species like SO₂, alkali chlorides, and HF. That completely changes the failure hierarchy. HRSG problems tend to cluster around flow-accelerated corrosion in economizer circuits and duct burner control instability. WHRB problems are dominated by fouling, erosion, and dew-point corrosion. Applying HRSG maintenance logic to a cement WHRB is a reliable way to miss the real failure mode until something cracks.
How often should WHRB safety valves be tested and recertified?
Annual functional testing is the baseline — required under ASME Section I, PED 2014/68/EU, and GB/T 12241, whichever governs your installation. Recertification by an approved third-party inspection body every three years is standard practice. One point often skipped: always test after any pressure excursion event, even if the valve appears to have reseated cleanly. A valve that lifted under overpressure and reseated may have a compromised disc-to-seat contact. Skipping that post-event test is how a plant ends up with a relief device that looks fine on the tag and fails to open when it’s actually needed.
What water treatment is recommended for a 2.5 MPa WHRB on hard groundwater?
At 2.5 MPa, the water chemistry requirements are genuinely demanding. Ion exchange softening to below 0.03 mmol/L hardness is the baseline. Thermal deaeration targeting dissolved oxygen below 7 µg/L is non-negotiable — oxygen pitting at this pressure range is fast and localized. Continuous phosphate dosing maintains residual alkalinity in the drum. Where source water TDS exceeds roughly 500 mg/L, adding RO pretreatment upstream of the softener reduces blowdown frequency meaningfully and extends resin service intervals. Online conductivity monitoring at economizer inlet and steam drum blowdown points gives you real-time chemistry control. Skipping the RO stage on high-TDS supply water to save capital cost is one of the more predictable ways to end up with heavy scale and tube failures inside 18 months.
Taishan Group WHRB field service engineers can be dispatched within 72 hours for critical issues under an active service agreement.True
This is a stated service commitment applicable to clients with signed service agreements; response time for non-critical remote support via VPN-connected DCS access and video inspection guidance is available on a shorter timeline.
How does Taishan Group support overseas clients with WHRB troubleshooting remotely?
Remote support starts with VPN-linked DCS data access — reviewing actual trend data is far more useful than a client trying to describe symptoms over email. For visual issues like refractory damage or tube surface condition, structured video-guided inspection protocols let a factory technician walk the boiler systematically while a Taishan engineer directs the inspection in real time. Fault report templates standardize what data gets captured so nothing diagnostic gets missed. For critical faults — pressure boundary damage, uncontrolled steam loss, severe drum level instability — field service engineers can be dispatched within roughly 72 hours under an active service agreement, covering most major industrial regions in Southeast Asia, the Middle East, and Africa where Taishan has commissioned units.
References
ASME BPVC Section VII — Recommended Guidelines for the Care of Power Boilers
Source: American Society of Mechanical Engineers (ASME)Suggested Maintenance Log Program for Boiler Systems
Source: The National Board of Boiler and Pressure Vessel InspectorsRecommendations for Developing a Preventive Boiler Maintenance Schedule
Source: The National Board of Boiler and Pressure Vessel InspectorsBoiler Logs Can Reduce Accidents
Source: The National Board of Boiler and Pressure Vessel InspectorsHRSG Tests, Inspections and Other Site Services
Source: GE VernovaHRSG Upgrades for Pressure Parts and Non-Pressure Parts
Source: GE VernovaHRSG Anomaly Detection — Performance Loss and Early Leak Detection
Source: GE VernovaPressureWave+ — Improving Efficiency with HRSG Cleaning
Source: GE VernovaHRSG Fouling and Cleaning Case Study — Florida
Source: GE VernovaWater for the Boiler — Water Quality, Carryover and Boiler Operation
Source: Spirax Sarco
How to Troubleshoot Common Issues in Waste Heat Recovery Boilers? Read More »

