I wanted to spider a website and, if some text or a matching pattern is found in the HTML, get the URL(s) of the page(s).
Wrote the command
wget --recursive --spider site.com 2>&1 | sort | uniq | grep -oe 'www[^ ]*'
to get all the URLs thus far, but stuck as to how to output only those URLs that have a specified text. Any clues?