TLDRs newsletters are TLDR so I wrote a converter that gets all the important links
Read the original on christianheilmann.com ↗Much like a lot of other people, I subscribe to quite a few of the newsletters of TLDR. As I re-use the things I found there in own publications, I did find the newsletters – ironically – not quite succinct enough to go through them quickly. I wanted to skip sponsored sections and other cruft, so, of course I wrote a PHP script to help me with that. Here it is and you can find it on GitHub.
$URLS = file_get_contents(‘tldrs.txt’);
$out = ‘’;
foreach (array_filter(explode(“n”, $URLS)) as $url) {
echo $url . “n”;
$html = file_get_contents($url);
$dom = new DOMDocument();
libxml_use_internal_errors(true);
$dom->loadHTML($html);
libxml_clear_errors();
$xpath = new DOMXPath($dom);
$links = $xpath->query(‘//a’);
foreach($links as $link) {
$content = $link->nodeValue;
if (strstr($content,’minute read’) ||
strstr($content,’GitHub Repo’)) {
$text = trim(
preg_replace(
“/ (d+ minute read)|(GitHub Repo)/”,
“”,$link->nodeValue)
);
$details = trim($link->parentNode->lastChild->previousSibling->nodeValue);
$link = $link->getAttribute(‘href’);
if (strpos($link, ‘links.tldrnewsletter.com’) !== 0) {
$link = exec(‘curl -Ls -o /dev/null -w %{url_effective} ’ . escapeshellarg($link));
}
$out .= trim($text).” ” . $link . “n” . $details . “nn”;
}
}
}
file_put_contents(‘tldrs.txt’, $URLS . “n” . $out);
?>
The script takes a text file called “tldrs.txt” and loads the URLs on each of its lines. For example:
https://a.tldrnewsletter.com/web-version?ep=1&lc=12287b12-8315-11ee-beaa-736450d0f676&p=f7d1a838-be26-11f1-8893-bde744e2a96e&pt=campaign&t=1790939613&s=ca95811254756b8f90a40fb908f378eb263ee5efd26c4f4e75cf27afe12c4c75
https://a.tldrnewsletter.com/web-version?ep=1&lc=072c9ebe-8315-11ee-bfc9-3105a9c4cdee&p=a87ac9ce-be42-11f1-81e5-6329ebe08e36&pt=campaign&t=1790937594&s=593f4a56fb8e145a55f12d69fbe7ad37ae0b8f79db13cdb50873af64859a8d24
https://a.tldrnewsletter.com/web-version?ep=1&lc=122aee56-8315-11ee-88bc-f3604afdd80b&p=dacd0d30-bd87-11f1-bb10-f5a7490d358d&pt=campaign&t=1790861448&s=180e6f7bff6862c042d890cffd0d0b86c48161084c379294691dd212ee59c658
https://a.tldrnewsletter.com/web-version?ep=1&lc=122aee56-8315-11ee-88bc-f3604afdd80b&p=102bee2e-be55-11f1-a985-65716d1d02e4&pt=campaign&t=1790947629&s=6f9c6a5ccbd4e6b92b9ac81680b0f4a037611ae42219f5859937387a97f21526
It loads the URLs one by one, and checks which links contain either `minutes read` or `GitHub Repo`. It then loads these and extracts their URL, headline and description and adds the results to the `tldrs.txt` file.
Following redirects
Some of the links in the newsletters are hidden behind a redirect of `links.tldrnewsletter.com`. As I wanted the real URLs, I am using cURL to get the correct ones.
You can do this with the following command:
curl Ls -o /dev/null -w %{url_effective} https://linkto-unfold
Pretty sure noone but me needs this, but here we are.