In the time of APIs, python scrapers and other cool things, bookmarklets have fallen from favor but they still have uses and I still like them. For small jobs that you want to make accessible to people they are still really handy.
In this scenario Middlebury has lots of ways to get roster information about students but for this particular scenario we needed the student first name, last name, and their email address. We couldn’t get a nice export where all three of those things existed. We did have a web page where it existed but there was no export option. We could have cut/pasted but in this scenario we had a bunch of sections so that would be a hassle.
I thought that a Python scraper was overdoing things and thought I could manage it with a little javascript bookmarklet.
So what I did to work this out was to cut/paste a chunk of the roster information into a page of my own and start working out the javascript portion. Could I find the pieces and identify them?
Below you can see a simplified version. I copied the HTML for one person and started playing with getting the data into separate arrays. It’s embedded below in CodePen if you want to see it all together but I’ll break down the key pieces below.
See the Pen
scraper demo by Tom (@twwoodward)
on CodePen.
Names
const names = document.querySelectorAll('h3');
names.forEach(function(name){
//console.log(name.innerText)
let split = name.innerText.split(", ");
first.push(split[1]);
last.push(split[0]);
// console.log(first)
// console.log(last)
})
In the HTML all of the student names are in H3 tags and they are the only H3 tags on the page. That’s awesome. It means we can use querySelectorAll(‘h3’) and grab them.
Now that we have them, we can loop through and split them up. In this case the names are given last name first and they’re separated by a comma and a space.
So our loop, does two things. It splits the innerText of the H3 tag by ‘, ‘. That returns an array with two items. The first item (0) is our last name and the second item (1) is our first name. Then it pushes the result into the arrays we have to hold the first and last names respectively.
You can see the console.log pieces I have in there that are commented out. My advice to anyone starting to code is to use these types of tools early and often. It’s the difference between playing tag and playing tag blindfolded.
This one was a bit harder. There are multiple email addresses in there for deans and stuff. The student email ends up being in there twice. You can see the student email HTML block below. Kind of weird but in this case we can use it to our advantage.
<dd class="d-inline-block">
<a href="mailto:fakename@middlebury.edu" class="link-underline"> </a><a href="mailto:fakename@middlebury.edu">fakename@middlebury.edu</a>
</dd>
[javscript]
const links = document.querySelectorAll(‘a’)
var email = [];
links.forEach(function(link){
//console.log(link.innerText)
if(link.innerText == ” && link.classList.contains(‘link-underline’)){
let cleanLink = link.getAttribute(“href”).substring(7);
email.push(cleanLink)
// console.log(email)
}
})
return email;
[/javascript]
Here we are going to grab all the ‘a’ tags. We know we’ll get the email in there somewhere but we’ll also get other links and emails we don’t want. That’s fine.
To filter them out we set two conditions, first the innerText needs to be empty (that odd structure is now our friend) and second that the link has to have the class ‘link-underline.’ These were good enough to get me just the student emails.
It can get a lot harder (or easier) with other patterns. You might have to select the third child after a particular element . . . stuff like that . . . but if there’s a consistent pattern then you can do it.
To CSV
I did this before and was able to find the same Stackoverflow response. So nice for people to do this kind of thing.
To the bookmarklet
I always mess up trying to get the javascript into the bookmarklet. That’s were Peter Coles and this page comes into play. I do the include custom script option as that’s the easiest for me. I can then just edit my javascript as normal without having to constantly reformat it to run in the bookmarklet. You could probably set up gulp or something to deal with this but why bother if you don’t have to?
I end up with this js. I can add that as the URL in an A HREF and now people can drag/drop it righ to their own bookmark bar.
javascript:(function()%7Bfunction callback()%7Bwindow.getNames()%7Dvar s%3Ddocument.createElement("script")%3Bs.src%3D"https://experiments.middcreate.net/extras/scraper/stu-data.js"%3Bif(s.addEventListener)%7Bs.addEventListener("load",callback,false)%7Delse if(s.readyState)%7Bs.onreadystatechange%3Dcallback%7Ddocument.body.appendChild(s)%3B%7D)()
I love me some nifty bookmarklet action. I use mine several times a day.
Do you run into CORS issues getting access to the stuff you are scraping? That’s been a roadblock for some of mine
It did when I tried to run from Gist initially . . . CORB I think.
I’d be curious about the issues though. I thought for stuff like this it’d be basically running in the browser like I entered it in console but I don’t actually know that.