[HN Gopher] Tell HN: Replace the X with a 5 in arXiv.org to disp...
       ___________________________________________________________________
        
       Tell HN: Replace the X with a 5 in arXiv.org to display a paper in
       HTML
        
       Check it out: https://ar5iv.org/pdf/2106.10522.pdf
        
       Author : bgschulman31
       Score  : 300 points
       Date   : 2022-02-01 15:00 UTC (8 hours ago)
        
       | chubot wrote:
       | Very nice idea, although the first PDF I tried wasn't available,
       | and the second one seems to be missing a lot:
       | 
       | https://arxiv.org/pdf/2001.00888.pdf (19 pages)
       | 
       | https://ar5iv.org/html/2001.00888 (missing content)
       | 
       | This is no doubt a hard problem ...
        
         | tech-no-logical wrote:
         | seems it craps out on the monospaced bits of text ?
        
         | dginev wrote:
         | If you have a spare minute, please pay a visit to the "report
         | issue" button at the bottom.
         | 
         | Indeed - hard problem and a messy solution. I have no easy
         | answers.
        
       | beanjuiceII wrote:
       | wow thank you for this!
        
       | steelstraw wrote:
       | Thanks!
       | 
       | What other sites have these kind of URL hacks?
        
         | slig wrote:
         | Youtube has the type "nsfw" before any youtube.com address to
         | bypass the age-restriction.
        
         | galori wrote:
         | replace reddit.com with redditp.com to make it into a slide
         | show of the images in the sub
         | 
         | example:
         | 
         | https://www.reddit.com/r/gifs/ -->
         | https://www.redditp.com/r/gifs/
         | 
         | (there are a few other similar reddit --> image gallery url
         | hacks, just google for them)
        
       | dark-star wrote:
       | I checked the example link, and it has trouble rendering utf-8 or
       | something (search for the Feynman quote about "nature", it starts
       | with "Nature isnat classical...", or the poem at the end)
       | 
       | I wonder how well it works with more complicated mathematical
       | formulas containing greek or arabic letters (although the example
       | in that paper look fine to me), or other non-ASCII scripts.
       | 
       | other than that, this looks pretty amazing
        
       | aasasd wrote:
       | By the way: if you can, it would probably be wise to drop
       | justified alignment, and make text aligned to the left. Rectangle
       | blocks of text only look good from a distance--when actually
       | trying to read, the differing inter-word spaces just make the
       | experience jarring. It's especially bad on phones, which is a
       | prime use-case for the HTML conversion.
       | 
       | (Though, from 'tell HN', the poster is probably not the site
       | author, right?)
        
         | dginev wrote:
         | HN can obviously summon the site author (hi!)
         | 
         | What I was wondering - and still am - shouldn't it be possible
         | to get to a "Pleasant" justified layout on the web in general?
         | 
         | The jagged left-aligned paragraphs are some of the first bits
         | people point to when invoking "my PDF looks better". I
         | definitely am not saying I did it perfectly, but shouldn't it
         | be _possible_ to get a good justified scientific article on the
         | web? Why not?
        
           | aasasd wrote:
           | Hi.
           | 
           | "Looks better" only works in regard to justification when one
           | is admiring a page overall. However, that's not how people
           | actually read text. They look at words in lines, and at that
           | time uneven spacing keeps tripping the eye up. I know this
           | argument, but there's no way around this discussion, and
           | that's all there is to say about it (so far). Nicely looking
           | bricks of paragraphs won't make the eye glide smoothly over
           | the holes.
           | 
           | Page layout programs spend some CPU time on fiddling the
           | hyphenation until the spacing is even. Browsers can't afford
           | to do that, afaik (not sure about currently, but that was the
           | situation a while back). Moreover, from what I vaguely heard,
           | the HTML specification defines paragraphs in such a way that
           | browsers don't even have the freedom to fiddle the paragraph
           | height--something about the height being the minimum for the
           | text on hand, or something like that.
           | 
           | Even if you insist to keep justification on desktop (though I
           | personally can see the holes clearly)--for the love of good,
           | please disable it on phones. It's just a mess there.
        
             | dginev wrote:
             | Got it. So even with hyphenation, the gaps are still bad
             | enough that you'd consider the current ar5iv rendering
             | bumpy and distracting?
             | 
             | I think I can see that, but it's _almost_ there, which is
             | why it feels like there has to be something I 'm missing
             | for it to justify "just right".
             | 
             | But yes, at the least you've convinced me we should have a
             | separate theme that goes left-aligned, and possibly makes a
             | number of other choices that maximize readability.
             | 
             | Since I'd still want the folks that want "as good as PDF",
             | to feel justified for sticking around.
        
               | aasasd wrote:
               | I _may_ have a better eye for it than most, all the way
               | to having made my own browser extension to turn
               | justification on paragraphs off. However, if I do turn on
               | left-aligning on the linked example, I get plenty of very
               | jagged right edges--which means that with justification
               | all that uneven space gets shoved between words.
        
       | hk1337 wrote:
       | This must be a Dartmouth project.
        
         | dginev wrote:
         | What's a Dartmouth?
        
       | hyperhopper wrote:
       | My biggest complaint with all these super useful sites is that I
       | can never remember them when I need them. Replace this in YouTube
       | to bypass country restrictions, replace that in arxiv to view in
       | browser, etc.
       | 
       | I wish somebody could make an extension or repository system to
       | store all these, and prompt you sometimes when on the sites.
        
         | sundarurfriend wrote:
         | Ah, this comment prompted me to add this to my Anki, so thanks!
         | I got frustrated never remembering https://remove-js.com , so
         | added it to a deck on Anki and now I'm unlikely to ever forget
         | it. This is useful enough to go on the deck too.
        
           | siva7 wrote:
           | I'm using Anki to not forget my fitness workouts
        
             | kazinator wrote:
             | Anki is great for remembering your fitness workouts, if you
             | plan to sabotage your fitness with increasing intervals
             | between workouts. :)
        
               | Kelamir wrote:
               | Hihi. That's true. Maybe an addon that keeps cards' ease
               | at the same rate would do.
               | 
               | By the way, have you checked out Migaku's vacation add-
               | on? I suggested it to you a while ago.
        
           | theWreckluse wrote:
           | Ah, I always forget about Anki! Wish there were an app to
           | remind me of Anki when I need it the most.
        
             | siva7 wrote:
             | The anki app reminds you of anki
        
           | gcoladon wrote:
           | I love Anki and use it like you do, it sounds like.
        
           | xigoi wrote:
           | Does RemoveJS do anything that the "block scripts" button on
           | uBlock doesn't?
        
             | sundarurfriend wrote:
             | Allows me to share it as a link to my parents, for example.
             | 
             | During the initial pandemic period, there were news
             | articles I wanted to share with useful information. But the
             | pages were full of Javascript based crap that guiding them
             | on how to find the information became a task of its own.
             | RemoveJS was very helpful in making those sites accessible
             | to them.
        
             | sli wrote:
             | I'd tell you, but RemoveJS blocks VPN users from viewing
             | their site while uBlock does not.
        
         | curiousllama wrote:
         | I agree, I can't wait to forget about what the name of the
         | repository is
        
           | EGreg wrote:
           | I can't wait to turn down at least $6B offer for my company
           | :)
        
         | FrozenVoid wrote:
         | I usually make a userscript for these sites, which add it
         | automatically. Its a one-liner for simple link insertion.
         | https://github.com/FrozenVoid/Userscripts/blob/main/Arxiv/Ar...
        
           | kazinator wrote:
           | The Redirector add-on for Firefox provides pattern-based
           | (regex and glob) URL redirection on the browser side. You can
           | just add rules to its configuration rather than new
           | userscripts.
           | 
           | I used in the past to locally correct for broken links within
           | intranet/CI pages. E.g. something correctably wrong with a
           | gerrit link, or whatever.
        
         | umvi wrote:
         | Replace "github.com" with "github.dev" or "github1s.com" on any
         | repo
        
       | brummm wrote:
       | The Latex created PDF's are 100 times more pleasant to read
       | though.
        
         | phailhaus wrote:
         | Problem is, the Latex created PDF's have fixed width and read
         | horribly on smaller screens as a result, with no option of a
         | dark mode. Often on mobile you'll be forced to download the PDF
         | as well.
        
       | baby wrote:
       | Do this with eprint please!
        
         | max_ wrote:
         | The entire project is open source[0].
         | 
         | [0]: https://github.com/dginev/ar5iv
        
       | fxtentacle wrote:
       | How does this compare to arxiv-vanity.org?
        
         | dginev wrote:
         | Same HTML backend generator (latexml), different frontends, and
         | different coverage of arXiv.
         | 
         | Also, ar5iv may disappear very quickly, since I am unsure if
         | it's more helpful or harmful. But I'll definitely lean on the
         | public attention to keep asking arXiv to integrate an HTML
         | preview for their articles. In the one-and-only arxiv.org
         | itself.
         | 
         | Lastly, one difference that may ignite a curious debate is that
         | ar5iv is committed to being MathML-native. Yes. MathML is the
         | only markup used for math syntax, and you'll see it rendered
         | directly, undisturbed, with Firefox today.
         | 
         | Over 500 million MathML elements in the full dataset too,
         | pretty awe-inspiring.
        
       | jppope wrote:
       | thats RAD. awesome work
        
       | swframe2 wrote:
       | Please consider adding a little bit of javascript to find and
       | display the text that describes the variable in a formula when
       | the mouse hovers over it. Even better would be to allow signed-in
       | users (or paid subscribers) to add annotations (and links to
       | youtube videos).
       | 
       | I would prefer if your site requires a paid subscription so you
       | can incentivize people to annotate the content. For a paper
       | author, making the paper terse and complex is more impressive but
       | for the rest of us it is very tedious to decipher.
        
       | resoluteteeth wrote:
       | Nice. It looks like a site with a similar idea was posted to hn
       | before[1], but the result from ar5iv seems to be a bit
       | slicker/cleaner than that site.
       | 
       | 1: https://www.arxiv-vanity.com/
        
         | jessriedel wrote:
         | Unfortunatley, ar5iv also only hosts the _first_ version of the
         | paper, while arxiv-vanity only hosts the _last_.
        
           | dginev wrote:
           | Which is intentional. ar5iv does not aim to be a live preview
           | service, or replace arXiv.
           | 
           | The primary aim is to serve the community with the outputs we
           | have, while we improve the coverage and fidelity of our
           | generator.
           | 
           | And yes, using only the official sources arXiv has released
           | for reuse: https://arxiv.org/help/bulk_data_s3
           | 
           | This is indeed a major difference with -vanity
        
             | sundarurfriend wrote:
             | > The primary aim is to serve the community with the
             | outputs we have, while we improve the coverage and fidelity
             | of our generator.
             | 
             | Could you explain what you mean here? Who is "we" - I
             | assume ar5iv? "while we improve the coverage and fidelity
             | of our generator" <-- does this mean this is a temporary
             | situation, and in the future multiple versions of the paper
             | will be available?
        
               | dginev wrote:
               | Certainly, sorry for the confusion.
               | 
               | There's actually multiple "we", since there are two
               | institutions involved, and one foil character - I'm the
               | only one responsible for ar5iv "the website", in a
               | personal capacity.
               | 
               | The fidelity of the generator has the "we" of the team
               | behind LaTeXML, the TeX-to-HTML conversion tool. That is
               | in many ways the most important project to remember here,
               | as that is what we want to actively improve to a point
               | where it is "good enough" in creating HTML over the
               | entirety of arXiv.
               | 
               | The institution hosting the website, and wanting to
               | "serve a community" is KWARC, a research group at the
               | university of FAU-Erlangen in Germany. There are all
               | kinds of projects and services brewing on that end, which
               | have interplay with the HTML data behind ar5iv, but are
               | not directly on the site.
               | 
               | And as to all of us reading HN, I think we are actually
               | interested in arXiv itself being maximally useful. And so
               | is the ar5iv site - it's a temporary deployment, that
               | really is aiming to reintegrate back into the arxiv.org
               | site, and general infrastructure.
               | 
               | If/when that happens is unclear, but in the meantime
               | there is a lot of improvements that can be made, both in
               | what HTML can be generated, deciding what the markup of
               | scientific documents _ought to be_ in the first place, as
               | well as gaining some insights for what new problems arXiv
               | would encounter if they served HTML.
        
               | dginev wrote:
               | Oh, and the last question - yes, if arXiv integrates the
               | feature, they will be able to serve any of the versions,
               | including the most recent one.
               | 
               | I can _technically_ implement that, but I really don 't
               | want to, as I see it as crossing a certain line. Seeing
               | ar5iv as a limited, constrained, service is a good thing
               | - I think it clearly communicates that _I_ do not want to
               | compete with arXiv.
        
         | visarga wrote:
         | Unlike Ar5iv, Arxiv-vanity will also show papers more recent
         | than a month, which are usually the papers you want to open.
        
       | gspr wrote:
       | That's a lot less bad than I expected!
        
         | dginev wrote:
         | Thanks!
        
       | sandhiw wrote:
       | Awesome, easy way to read papers in the browser with dark mode
       | without inverting image colors (like what Dark Reader does)
        
       | jvanderbot wrote:
       | Before you get too excited:
       | 
       | "We are usually at least a month behind the live arXiv article
       | list.
       | 
       | Also, we can only serve papers submitted with their LaTeX
       | sources."
       | 
       | Moderate excitement is probably warranted, though.
        
         | dynm wrote:
         | If a paper is prepared in latex, arxiv will detect that and
         | insist that the sources are included. (They compile the pdf
         | themselves.) So I think the second point shouldn't be much of a
         | problem in fields where everyone uses latex.
        
       | max_ wrote:
       | The project is open source, so this can easily be ported to any
       | site[0]. [0]: https://github.com/dginev/ar5iv
        
       | tingletech wrote:
       | arXiv have the advantage that they have LaTeX source for a
       | majority of their submissions. Much easier to convert that to
       | HTML that any arbitrary PDF.
        
         | evanb wrote:
         | For any arXiv submission that was submitted via LaTeX, you can
         | download the source.
        
         | jessriedel wrote:
         | This was a very far-sighted move by, I believe, Paul Ginsparg
         | back in the early days. It significantly increases the headache
         | of submitting to the arXiv because you have to get your tex
         | file to compile with the tex distribution on _their_ servers
         | rather than just on your own home box. But it makes the arxiv
         | vastly more future proof than it would be if you could just
         | upload PDFs.
        
       | scythmic_waves wrote:
       | This is great! The hoverable footnotes are a nice touch.
        
       | fullam wrote:
       | Need this for bioarxiv also!
        
         | michaelhoffman wrote:
         | bioRxiv converts PDFs to full-text already. They're right there
         | on the bioRxiv web site--just click the "Full text" tab. For
         | example, our most recent bioRxiv preprint:
         | 
         | https://www.biorxiv.org/content/10.1101/2022.01.07.475366v1....
         | 
         | If you want to do URL munging instead of using the UI, just add
         | `.full` to the end of the URL ;)
         | 
         | This doesn't require that the manuscript be submitted in any
         | particular format either. There are humans involved in making
         | sure the full text formatting from PDF is good.
        
       | aasasd wrote:
       | People put in effort making PDFs and ensuring that nobody can
       | read them without either a 14" portrait screen or a ton of
       | scrolling--and you had to come along and ruin that carefully laid
       | out inconvenience? What's wrong with you.
       | 
       | The incredibly complex layout of 'a wall of text and occasional
       | pictures' is there for a reason. That's what the authors wanted,
       | and only PDF is up to the task of representing such delicate
       | formatting.
        
         | [deleted]
        
         | shishy wrote:
         | Not sure if this is sarcasm... but I don't think OP made this
         | (https://github.com/dginev/ar5iv), they are just sharing.
         | 
         | Besides I don't think this is meant to be a total replacement
         | that everyone is going to magically use and abandon PDFs, it
         | might just be a convenience within a workflow to skim and check
         | the full PDF if necessary.
        
           | ouid wrote:
           | I think the holes on your sarcasm filter might be too large.
        
             | jimhefferon wrote:
             | I used to think I could recognize sarcasm whenever I saw
             | it. Then I lived through the past few years.
        
               | ouid wrote:
               | I'm not making detailed arguments about what your sarcasm
               | filter should be, but if this comment isn't detected as
               | sarcasm, you should turn a knob somewhere.
        
       | 3pt14159 wrote:
       | Thank you! This is amazing! My only ask is that you bump up the
       | default font size.
        
         | Symbiote wrote:
         | Don't most browsers remember your zoom when you revisit a
         | website?
        
           | 3pt14159 wrote:
           | Yes, but why have such an hard to read default?
        
             | Kelamir wrote:
             | I'm not sure but it could be that it makes it easier to
             | read on big displays. Of course this isn't very Responsive
             | Design like. I will be glad to be proven wrong. Interested
             | in why HN has such a small font by default as well.
        
           | eatonphil wrote:
           | Maybe it's a mobile issue? I've never adjusted zoom in my
           | phone.
        
         | kzrdude wrote:
         | I was going to suggest bump up the saturation/blackness of the
         | font so that it's easier to read.
        
       ___________________________________________________________________
       (page generated 2022-02-01 23:01 UTC)