Why doesn't Node.js have a native DOM?

Viewed 42701

When I discovered that Node.js was built using the V8 JavaScript engine, I thought:

Great, web scraping will be easier as the page will be rendered like in the browser, with a "native" DOM supporting XPath and any AJAX calls on the page executed.

  1. Why doesn't it have a native DOM when it uses the same JavaScript engine as Chrome?
  2. Why doesn't it have a mode to run JavaScript in retrieved pages?
  3. What am I not understanding about JavaScript engines vs the engine in a web browser?

Many thanks!

13 Answers

The Document Object Model (DOM in short) is a programming interface for HTML and XML documents and it represents the page so that programs can change the document structure, style, and content. More on this subject.


The necessary distinction between client-side (browser) and server-side (Node.js) and their main goals:

  • Client-side: accessing and displaying information of the web
  • Server-side: providing stable and reliable ways to deliver web information

Why is there no DOM in Node.js be default?

By default, Node.js doesn't have access, nor have any knowledge about the actual DOM in your own browser. Node.js just delivers the data, that will be used by your own browser to process and render the whole website, the DOM included. The server provides the data to your browser to use and process. That is the intended way.

Why wouldn't you want to access the DOM in Node.js?

Accessing your browser's actual DOM using Node.js would be just simply out of the goal of the server. Your own browser's role is to display the data coming from the server. However it is certainly possible and there are multiple solutions in different level of depths and varieties to pre-render, manipulate or change the DOM using AJAX calls. We'll see what future trends will bring.

Why would you want to access the DOM in Node.js?

By default, you shouldn't access your own, actual DOM (at least some data of it) using Node.js. Client-side and server-side are separated in terms of role, functionality, and responsibility based on years of experience and knowledge. Although there are several situations, where there are solid reasons to do so:

  • Gathering usage data (A/B testing, UI/UX efficiency and feedback)
  • Headless testing (Development, automation, web-scraping)

How can you access the DOM in Node.js?

  • jsdom: pure-JavaScript implementation, good for testing your own DOM/browser-related project
  • cheerio: great solution if you like/often use jQuery
  • puppeteer: Google's own way to provide headless testing using Google Chrome
  • own solution (your possible future project link here)

Although these solutions do not provide a way to access your browser's own, actual DOM by default, but you can create a project to send some form of data about your DOM to the server, then use/render/manipulate that data based on your needs.

...and yes, web-scraping and web development in terms of tools and utilities became more sophisticated and certainly easier in several fields.

node.js chose not to include it in their standard library. For any functionality, there is an inevitable tradeoff between comprehensiveness, scalability, and maintainability.

That doesn't mean it's not potentially useful. There is at least one JavaScript DOM implementation intended for NodeJS (among other CommonJS implementations).

2018 answer: mainly for historical reasons, but this may change in future.

Historically, very little DOM manipulation was done on the server. Addiotinally, as other answers allude, the JS stdlib and the DOM are seperate libraries - if you're using node, for, say, Unix scripting, then HTMLElement and NodeList etc aren't really relevant to that.

However: server-side DOM manipulation is now a very common part of delivering web apps. Web servers need to understand the structure of pages, and, if asked to render a resource as HTML, deliver HTML content that reflects the initial state of a web application. This means web apps load much faster than if the server simply delivers a stub page and has the browsers then do the work of filling in the real content. Currently this is done with JSDom and similar, but in the same way node has Request and Response objects built in, having DOM functions maintained as part of the stdlib would help with this task.

Node is a runtime environment, it does not render a DOM like a browser.

Because there isn't a DOM. DOM stands for Document Object Model. There is no document in Node, so not DOM to manipulate it. That is definitively a browser thing.

You can use a library like cheerio though which gives you some simple DOM manipulation.

Node is server-level JavaScript. It's just the language applied to a basic system API, more like C++ or Java.

Related