How do I cut HTML so that the closing tags are preserved?

Viewed 468

How can I create a preview of a blog post stored in HTML? In other words, how can I "cut" HTML, making sure the tags close properly? Currently, I'm rendering the whole thing on the frontend (with react's dangerouslySetInnerHTML) then setting overflow: hidden and height: 150px. I would much prefer a way where I could cut the HTML directly. This way I don't need to send the entire stream of HTML to the frontend; if I had 10 blog post previews, that would be a lot of HTML sent that the visitor would not even see.

If I had the HTML (say this was the entire blog post)

<body>
   <h1>Test</h1>
   <p>This is a long string of text that I may want to cut.. blah blah blah foo bar bar foo bar bar</p>
</body>

Trying to slice it (to make a preview) wouldn't work because the tags would become unmatched:

<body>
   <h1>Test</h1>
   <p>This is a long string of text <!-- Oops! unclosed tags -->

Really what I want is this:

<body>
   <h1>Test</h1>
   <p>This is a long string of text</p>
</body>

I'm using next.js, so any node.js solution should work fine. Is there a way I can do this (e.g. a library on the next.js server-side)? Or will I just have to parse the HTML myself (server-side) and then fix the unclosed tags?

2 Answers

post-preview


It was a challenging task and made me struggle for about two days and made me publish my first NPM package post-preview which can solve your problem. Everything is described in its readme, but if you want to know how to use it for your specific problem:

First of all install the package using NPM or download its source code from GitHub

Then you can use it before the user posts their blogpost to the server and send its result (preview) with the full post to the backend and validate its length and sanitize its html and save it to your backend storage (DB etc.) and send it back to users when you want to show them a blog post preview instead of the full post.

example:

The following code will accept the .blogPostContainer HTMLElement as input and returns the summarized HTML string version of it with *maximum 200 characters length.

You can see the preview in the 'previewContainer' .preview:

js:

import  postPreview  from  "post-preview";
const  postContainer = document.querySelector(".blogPostContainer");
const  previewContainer = document.querySelector(".preview");
previewContainer.innerHTML = postPreview(postContainer, 200);

html (complete blog post):

<div class="blogPostContainer">
  <div>
    <h2>Lorem ipsum</h2>
    <p>
      Lorem ipsum, dolor sit amet consectetur adipisicing elit. Neque, fugit hic! Quas similique
      cupiditate illum vitae eligendi harum. Magnam quam ex dolor nihil natus dolore voluptates
      accusantium. Reprehenderit, explicabo blanditiis?
    </p>
  </div>
  <p>
    Lorem ipsum dolor sit amet consectetur adipisicing elit. Ipsam non incidunt, corporis debitis
    ducimus eum iure sed ab. Impedit, doloribus! Quos accusamus eos, incidunt enim amet maiores
    doloribus placeat explicabo.Eaque dolores tempore, quia temporibus placeat, consequuntur hic
    ullam quasi rem eveniet cupiditate est aliquam nisi aut suscipit fugit maiores ad neque sunt
    atque explicabo unde! Explicabo quae quia voluptatem.
  </p>
</div>

<div class="preview"></div>

result (blog post preview):

<div class="preview">
  <div class="blogPostContainer">
    <div>
      <h2>Lorem ipsum</h2>
      <p>
        Lorem ipsum, dolor sit amet consectetur adipisicing elit. Neque, fugit hic! Quas similique
        cupiditate illum vitae eligendi ha
      </p>
    </div>
  </div>
</div>

It's a synchronous task so if you want to run it against multiple posts at once, you've better run it in a worker for better performance.

Thank you for making me do some research!

Good luck!

It is pretty complicated to guess what is the height of each pre-rendered element. However, you can cut the entry by number of characters with this pseudo rules:

    1. First define the maximum characters you want to keep.
    1. From the start: If you are meeting an HTML tag (Identify it by regexing < .. > or < .. />) go and find the closing tag.
    1. Then continue from where you stopped to search the tag.

A fast suggestion In javascript that I just wrote (probably can be improved, but that's the idea):

let str = `<body>
   <h1>Test</h1>
   <p>This is a long string of text that I may want to cut.. blah blah blah foo bar bar foo bar bar</p>
</body>`;

const MAXIMUM = 100; // Maximum characters for the preview
let currentChars = 0; // Will hold how many characters we kept until now

let list = str.split(/(<\/?[A-Za-z0-9]*>)/g); // split by tags

const isATag = (s) => (s[0] === '<'); // Returns true if it is a tag
const tagName = (s) => (s.replace('<', '').replace('>', '').replace('\/', '')) // Get the tag name
const findMatchingTag = (list, i) => {
    let name = tagName(list[i]);
    let searchingregex = new RegExp(`<\/ *${name} *>`,'g'); // The regex for closing mathing tag
    let sametagregex = new RegExp(`< *${name} *>`,'g'); // The regex for mathing tag (in case there are inner scoped same tags, we want to pass those)
    let buffer = 0; // Will count how many tags with the same name are in an inner hirarchy level, we need to pass those
    for(let j=i+1;j<list.length;j++){
        if(list[j].match(sametagregex)!=null) buffer++;
        if(list[j].match(searchingregex)!=null){
            if(buffer>0) buffer--;
            else{
                return j;
            }
        }
    }
    return -1;
}

let k = 0;
let endCut = false;
let cutArray = new Array(list.length);
while (currentChars < MAXIMUM && !endCut && k < list.length) { // As long we are still within the limit of characters and within the array
    if (isATag(list[k])) { // Handling tags, finding the matching tag
        let matchingTagindex = findMatchingTag(list, k);
        if (matchingTagindex != -1) {
            if (list[k].length + list[matchingTagindex].length + currentChars < MAXIMUM) { // If icluding both the tag and its closing exceeds the limit, do not include them and end the cut proccess
                currentChars += list[k].length + list[matchingTagindex].length;
                cutArray[k] = list[k];
                cutArray[matchingTagindex] = list[matchingTagindex];
            }
            else {
                endCut = true;
            }
        }
        else {
            if (list[k].length + currentChars < MAXIMUM) { // If icluding the tag exceeds the limit, do not include them and end the cut proccess
                currentChars += list[k].length;
                cutArray[k] = list[k];
            }
            else {
                endCut = true;
            }
        }
    }
    else { // In case it isn't a tag - trim the text
        let cutstr = list[k].substring(0, MAXIMUM - currentChars)
        currentChars += cutstr.length;
        cutArray[k] = cutstr;
    }
    k++;
}

console.log(cutArray.join(''))

Related