How to Remove Excess Whitespace in MODX + Fenom Without Joining Words Together

When you view the source code of a website built with MODX, pdoTools, and Fenom, you may sometimes notice a huge number of blank lines, spaces, and indentation.

The page itself usually looks fine in the browser, but the generated HTML can look like this:

<div class="author">

    
    
    <span>John</span>
    
    
    <span>William</span>
    
    
    <span>Smith</span>


</div>

In most cases, this extra whitespace does not affect how the page works. But on large websites with many nested chunks and templates, the final HTML source can become unnecessarily messy.

This becomes especially noticeable when a single page is assembled from dozens of Fenom chunks.

The obvious idea is to clean it up.

pdoTools provides a convenient option for exactly that:

{"strip": true}

At first, this seems like the perfect solution.

The generated HTML becomes much cleaner and more compact.

But in our case, it introduced a much more serious problem.

Why We Enabled strip in the First Place

Fenom allows templates to be written in a readable way:

<div class="author">
    {if $author}
        <span>{$author.lastName}</span>
        <span>{$author.firstName}</span>
        <span>{$author.middleName}</span>
    {/if}
</div>

This is easy for a developer to read and maintain.

However, line breaks and indentation from the template can also appear in the final HTML output.

The amount of whitespace grows even more when the template contains:

    nested chunks;
{if} conditions;
{foreach} loops;
{set} statements;
Fenom blocks;
snippet calls.

As a result, the page source may contain dozens of empty lines between actual elements.

From the browser's perspective, this is usually harmless.

But for production output, it is natural to want cleaner HTML.

So in the pdoTools system setting:

pdotools_fenom_options

we enabled:

{"strip": true}

At first, everything looked great.

The HTML became compact.

The Unexpected Problem: Words Started Joining Together

Later, a client noticed a strange issue.

For example, a template contained:

<span>John</span>
<span>William</span>
<span>Smith</span>

Before enabling strip, the browser displayed:

John William Smith

After enabling strip, it could become:

JohnWilliamSmith

The words were suddenly joined together.

At first glance, this seems odd because the text inside each <span> is still correct.

The problem is the whitespace between the tags.

Why a Line Break Between <span> Elements Becomes a Space

Consider ordinary HTML:

<span>John</span>
<span>William</span>

There is a line break between:

</span>

and:

<span>

In HTML, that line break is whitespace text.

The browser renders it as a normal space.

So visually we get:

John William

Now compare that with:

<span>John</span><span>William</span>

There is no text character between the two elements anymore.

The browser therefore displays:

JohnWilliam

This is exactly what happened after whitespace stripping.

Why We Could Not Simply Keep strip

The strip mechanism effectively transformed:

<span>John</span>
<span>William</span>

into:

<span>John</span><span>William</span>

For many HTML elements, this is completely safe.

For example:

</div>
<div>

can usually be converted to:

</div><div>

without changing the visual result.

That is because <div> is a block-level element.

Inline elements are different.

Examples include:

<span>
<strong>
<a>
<em>
<small>
<time>

These elements often appear inside normal text.

In that case, the whitespace between them may be part of the actual content, not just formatting.

Why Editing All Existing Content Was Not Practical

The most obvious workaround would be:

<span>John</span>&#32;<span>William</span>&#32;<span>Smith</span>

&#32; is the HTML entity for a regular space.

Even if line breaks are removed, this remains:

<span>John</span>&#32;<span>William</span>&#32;<span>Smith</span>

and the browser still displays:

John William Smith

Technically, this works.

Practically, it was not a reasonable solution.

The website already contained a large amount of:

  • pages;

  • articles;

  • chunks;

  • templates;

  • legacy HTML;

  • editor-generated content.

Manually searching the entire website for patterns like:

</span> <span>

and replacing whitespace with:

&#32;

would have been time-consuming and fragile.

It would also mean remembering to do the same thing every time new content was added.

That is not a good long-term architecture.

Templates should remain natural:

<span>John</span>
<span>William</span>

and the problem should be solved centrally.

Why We Could Not Just Configure Fenom strip

The next idea was to keep:

{"strip": true}

but customize its behavior.

For example:

Remove whitespace between div, section, and article, but preserve it between span elements.

Unfortunately, the standard Fenom strip behavior does not provide that level of control.

It is essentially enabled or disabled.

So instead of trying to change strip itself, we decided to intercept the Fenom template before the standard stripping logic ran.

For that, pdoTools provides the event:

pdoToolsOnFenomInit

What pdoToolsOnFenomInit Does

pdoTools uses Fenom to process templates.

Before Fenom starts compiling templates, pdoTools gives us access to the Fenom instance through:

pdoToolsOnFenomInit

At that point, we can register a custom Fenom filter.

The processing flow becomes:

Fenom template
     ↓
custom filter
     ↓
protect important whitespace
     ↓
standard strip
     ↓
compilation
     ↓
final HTML

The key idea is that we do not modify the final HTML after everything has already been rendered.

We protect important whitespace before the standard strip operation removes it.

Which Spaces Need to Be Protected

We are interested in patterns like:

</span>
<span>

or:

</strong> <span>

or:

</a>
<strong>

In other words:

closing inline tag
+
one or more whitespace characters
+
opening inline tag

For example:

</span>

    <span>

should become:

</span>&#32;<span>

After that, Fenom can still strip regular whitespace around it.

But it cannot remove &#32;, because this is no longer an actual whitespace character.

The browser later interprets &#32; as a normal space.

Which HTML Tags Are Treated as Inline

The plugin uses a list of inline elements:

$inlineTags = [
    'a',
    'abbr',
    'b',
    'bdi',
    'bdo',
    'cite',
    'code',
    'del',
    'dfn',
    'em',
    'i',
    'ins',
    'kbd',
    'mark',
    'q',
    's',
    'samp',
    'small',
    'span',
    'strong',
    'sub',
    'sup',
    'time',
    'u',
    'var',
];

The filter does not try to preserve whitespace between every possible pair of HTML elements.

For example:

</div>
<div>

does not need special protection.

But this does:

</span>
<span>

That makes the behavior much safer and more targeted.

How the Regular Expression Works

The central part of the solution looks like this:

preg_replace(
    '~(</(?:' . $inlineTags . ')\s*>)\s+(<(?:' . $inlineTags . ')\b)~iu',
    '$1&#32;$2',
    $text
);

It looks complicated at first, but the idea is simple.

The expression searches for:

</inline-tag>

then:

spaces / tabs / line breaks

and then:

<inline-tag

For example:

</span>
    <span>

is converted into:

</span>&#32;<span>

The contents of the tags themselves are not changed.

What the Result Looks Like

Original template:

<div class="author">
    <span>John</span>
    <span>William</span>
    <span>Smith</span>
</div>

After our custom filter:

<div class="author">
    <span>John</span>&#32;<span>William</span>&#32;<span>Smith</span>
</div>

After standard Fenom strip:

<div class="author"><span>John</span>&#32;<span>William</span>&#32;<span>Smith</span></div>

But the user still sees:

John William Smith

So we get both benefits:

  • compact HTML;

  • correctly separated words.

Why a Generic Whitespace Regex Is Dangerous

A common HTML minification trick is something like:

$html = preg_replace(
    '/>\s+</',
    '><',
    $html
);

At first glance, this looks perfect.

For example:

<div>

    <p>Text</p>

</div>

becomes:

<div><p>Text</p></div>

But the same rule also changes:

<span>Hello</span>
<span>world</span>

into:

<span>Hello</span><span>world</span>

and the browser renders:

Helloworld

So HTML minification is more complicated than simply:


Remove everything between > and <

Whitespace can be part of the visible content.

Why a Plugin Is Better Than Editing Templates

With the plugin in place, there was no need to modify:

  • existing articles;
  • old chunks;
  • templates;
  • TinyMCE content;
  • HTML stored inside MODX resources.

Editors and developers can continue writing normal HTML:

<span>First word</span>
<span>Second word</span>

The system itself protects the meaningful space by converting it to:

&#32;

before Fenom strips the rest.

This is especially useful on websites with hundreds or thousands of pages.

Instead of fixing the same problem in hundreds of places, the behavior is corrected once at the template engine level.

How to Create the Plugin in MODX

In the MODX Manager, create a regular plugin.

For example:

FenomSafeStrip

Attach it to the system event:

pdoToolsOnFenomInit

The pdoTools system setting:

pdotools_fenom_options

can remain:

{"strip": true}

The custom plugin protects important whitespace before Fenom runs its normal strip logic.

After creating or changing the plugin, clear the MODX cache.

Full Plugin Code

Below is the complete plugin implementation.

<?php

/**
 * FenomSafeStrip
 *
 * MODX / pdoTools / Fenom
 *
 * Event:
 * pdoToolsOnFenomInit
 *
 * Allows using:
 *
 * pdotools_fenom_options = {"strip": true}
 *
 * without joining words between adjacent
 * inline HTML elements.
 *
 * Example:
 *
 * <span>Smith</span>
 * <span>John</span>
 *
 * becomes before strip:
 *
 * <span>Smith</span>&#32;<span>John</span>
 */


/**
 * Get the Fenom instance.
 */
if (
    !isset($fenom)
    || !is_object($fenom)
) {
    return;
}


/**
 * Inline HTML elements where whitespace
 * may be part of visible text.
 */
$inlineTags = implode('|', [
    'a',
    'abbr',
    'b',
    'bdi',
    'bdo',
    'cite',
    'code',
    'del',
    'dfn',
    'em',
    'i',
    'ins',
    'kbd',
    'mark',
    'q',
    's',
    'samp',
    'small',
    'span',
    'strong',
    'sub',
    'sup',
    'time',
    'u',
    'var',
]);


/**
 * Add a Fenom pre-filter.
 *
 * It runs before standard AUTO_STRIP.
 */
$fenom->addFilter(
    static function ($template, $text) use ($inlineTags) {

        /**
         * Search for:
         *
         * </inline-tag>
         *     whitespace
         * <inline-tag>
         *
         * Example:
         *
         * </span>
         * <span>
         *
         * and replace that whitespace with:
         *
         * </span>&#32;<span>
         *
         * After that, strip can no longer remove
         * the meaningful space.
         */
        $result = preg_replace(
            '~(</(?:'
                . $inlineTags
                . ')\s*>)\s+(<(?:'
                . $inlineTags
                . ')\b)~iu',
            '$1&#32;$2',
            $text
        );


        /**
         * If preg_replace fails,
         * return the original template.
         */
        return $result ?? $text;
    }
);

How to Test It

Start with a simple test:

<p>
    <span>John</span>
    <span>William</span>
    <span>Smith</span>
</p>

With:

{"strip": true}

the resulting HTML source should look approximately like this:

<p><span>John</span>&#32;<span>William</span>&#32;<span>Smith</span></p>

But visually, the browser should still display:

John William Smith

You should also test combinations such as:

<strong>Very</strong>
<span>important</span>
<em>text</em>

and:

<a href="#">First</a>
<a href="#">Second</a>

At the same time, normal block-level elements should still be compacted:

</div>
<div>

Important Limitation

This plugin is not a universal HTML minifier.

It solves one specific problem:

It allows Fenom strip to be used while preserving meaningful whitespace between adjacent inline HTML elements.

That limited scope is actually a benefit.

The more aggressively we try to “optimize” HTML using broad regular expressions, the greater the chance of accidentally changing page content.

Conclusion

Sometimes an optimization that looks completely harmless can unexpectedly affect the actual content of a website.

The Fenom option:

{"strip": true}

does a good job of removing unnecessary indentation and line breaks.

But HTML whitespace is not always just formatting noise.

In this example:

<span>John</span>
<span>William</span>

the line break also acts as the visible space between words.

Removing it globally can turn:

John William

into:

JohnWilliam

Instead of modifying hundreds of existing pages, we moved the solution to the template engine level:

original Fenom template
        ↓
pdoToolsOnFenomInit
        ↓
protect whitespace between inline elements
        ↓
&#32;
        ↓
Fenom strip
        ↓
compact HTML
        ↓
text remains readable

The result is cleaner HTML without forcing editors or developers to manually add &#32; throughout existing content.

The broader lesson is useful beyond MODX:

when the same problem appears across hundreds of pages, it is usually better to fix the processing rule once than to edit hundreds of pages manually.

Do you want to start a project? Book a free call with me

Contact us
Cookie Policy
This website uses cookies to ensure you get the best experience on our website. By continuing to use this site, you consent to the use of cookies. Our Privacy Policy provides more information and explains how to amend your cookie settings.
Ok, accept
Ask my AI assistant