DOI-like strings and fake DOIs

TL;DR

Crossref discourages our members from using DOI-like strings or fake DOIs.

discouraged

Details

Recently we have seen quite a bit of debate around the use of so-called “fake-DOIs.” We have also been quoted as saying that we discourage the use of “fake DOIs” or “DOI-like strings”. This post outlines some of the cases in which we’ve seen fake DOIs used and why we recommend against doing so.

Using DOI-like strings as internal identifiers

Some of our members use DOI-like strings as internal identifiers for their manuscript tracking systems. These only get registered as real DOIs with Crossref once an article is published. This seems relatively harmless, except that, frequently, the unregistered DOI-like strings for unpublished (e.g. under review or rejected manuscripts) content ‘escape’ into the public as well. People attempting to use these DOI-like strings get understandably confused and angry when they don’t resolve or otherwise work as DOIs. After years of experiencing the frustration that these DOI-like things cause, we have taken to recommending that our members not use DOI-like strings as their internal identifiers.

Using DOI-like strings in access control compliance applications

We’ve also had members use DOI-like strings as the basis for systems that they use to detect and block tools designed to bypass the member’s access control system and bulk-download content. The methods employed by our members have fallen into two broad categories:

  • Spider (or robot) traps.
  • Proxy bait.

Spider traps

spider trap

A “spider trap” is essentially a tripwire that allows a site owner to detect when a spider/robot is crawling their site to download content. The technique involves embedding a special trigger URL in a public page on a web site. The URL is embedded such that a normal user should not be able see it or follow it, but an automated bot (aka “spider”) will detect it and follow it. The theory is that when one of these trap URLs is followed, the website owner can then conclude that the ip address from which it was followed harbours a bot and take action. Usually the action is to inform the organisation from which the bot is connecting and to ask them to block it. But sometimes triggering a spider trap has resulted in the IP address associated with it being instantly cut off. This, in turn, can affect an entire university’s access to said member’s content.

When a spider/bot trap includes a DOI-like string, then we have seen some particularly pernicious problems as they can trip-up legitimate tools and activities as well. For example, a bibliographic management browser plugin might automatically extract DOIs and retrieve metadata on pages visited by a researcher. If the plugin were to pick up one of these spider traps DOI-like strings, it might inadvertently trigger the researcher being blocked- or worse- the researcher’s entire university being blocked. In the past, this has even been a problem for Crossref itself. We periodically run tools to test DOI resolution and to ensure that our members are properly displaying DOIs, CrossMarks, and metadata as per their member obligations. We’ve occasionally been blocked when we ran across the spider traps as well.

Proxy bait

proxy bait

Using proxy bait is similar to using a spider trap, but it has an important difference. It does not involve embedding specially crafted DOI like strings on the member’s website itself. The DOI-like strings are instead fed directly to tools designed to subvert the member’s access control systems. These tools, in turn, use proxies on a subscriber’s network to retrieve the “bait” DOI-like string. When the member sees one of these special DOI-like strings being requested from a particular institution, they then know that said institution’s network harbours a proxy. In theory this technique never exposes the DOI-like strings to the public and automated tools should not be able to stumble upon them. However, recently one of our members had some of these DOI-like strings “escape” into the public and at least one of them was indexed by Google. The problem was compounded because people clicking on these DOI-like strings sometimes ended having their university’s IP address banned from the member’s web site. As you can imagine, there has been a lot of gnashing of teeth. We are convinced, in this case, that the member was doing their best to make sure the DOI-like strings never entered the public. But they did nonetheless. We think this just underscores how hard it is to ensure DOI-like strings remain private and why we recommend our members not use them.

Pedantry and terminology

Notice that we have not used the phrase “fake DOI” yet. This is because, internally, at least, we have distinguished between “DOI-like strings” and “fake DOIs.” The terminology might be daft, but it is what we’ve used in the past and some of our members at least will be familiar with it. We don’t expect anybody outside of Crossref to know this.

To us, the following is not a DOI:

10.5454/JPSv1i220161014

It is simply a string of alphanumeric characters that copy the DOI syntax. We call them “DOI-like strings.” It is not registered with any DOI registration agency and one cannot lookup metadata for it. If you try to “resolve” it, you will simply get an error. Here, you can try it. Don’t worry- clicking on it will not disable access for your university.

http://doi.org/10.5454/JPSv1i220161014

The following is what we have sometimes called a “fake DOI”

10.5555/12345678

It is registered with Crossref, resolves to a fake article in a fake journal called The Journal of Psychoceramics (the study of Cracked Pots) run by a fictitious author (Josiah Carberry) who has a fake ORCID (http://orcid.org/0000-0002-1825-0097) but who is affiliated with a real university (Brown University).

Again, you can try it.

http://doi.org/10.5555/12345678

And you can even look up metadata for it.

http://api.crossref.org/works/10.5555/12345678

Our dirty little secret is that this “fake DOI” was registered and is controlled by Crossref.

Why does this exist? Aren’t we subverting the scholarly record? Isn’t this awful? Aren’t we at the very least hypocrites? And how does a real university feel about having this fake author and journal associated with them?

Well- the DOI is using a prefix that we use for testing. It follows a long tradition of test identifiers starting with “5”. Fake phone numbers in the US start with “555”. Many credit card companies reserve fake numbers starting with “5”. For example, Mastercard’s are “5555555555554444” and “5105105105105100.”

We have created this fake DOI, the fake journal and the fake ORCID so that we can test our systems and demonstrate interoperable features and tools. The fake author, Josiah Carberry, is a long-running joke at Brown University. He even has a Wikipedia entry. There are also a lot of other DOIs under the test prefix “5555.”

We acknowledge that the term “fake DOI” might not be the best in this case- but it is a term we’ve used internally at least and it is worth distinguishing it from the case of DOI-like strings mentioned above.

But back to the important stuff….

As far as we know, none of our members has ever registered a “fake DOI” (as defined above) in order to detect and prevent the circumvention of their access control systems. If they had, we would consider it much more serious than the mere creation of DOI-like strings. The information associated with registered DOIs becomes part of the persistent scholarly citation record. Many, many third party systems and tools make use of our API and metadata including bibliographic management tools, TDM tools, CRIS systems, altmetrics services, etc. It would be a very bad thing if people started to worry that the legitimate use of registered DOIs could inadvertently block them from accessing content. Crossref DOIs are designed to encourage discovery and access- not block it.

And again, we have absolutely no evidence that any of our members has registered fake DOIs.

But just in case, we will continue to discourage our members from using DOI-like strings and/or registering fake DOIs.

This has been a public service announcement from the identifier dweebs at Crossref.

Image Credits

Unless otherwise noted, included images purchased from The Noun Project

Clinical trial data and articles linked for the first time

It’s here. After years of hard work and with a huge cast of characters involved, I am delighted to announce that you will now be able to instantly link to all published articles related to an individual clinical trial through the CrossMark dialogue box. Linked Clinical Trials are here!

In practice, this means that anyone reading an article will be able to pull a list of both clinical trials relating to that article and all other articles related to those clinical trials – be it the protocol, statistical analysis plan, results articles or others – all at the click of a button. Continue reading “Clinical trial data and articles linked for the first time”

Distributed Usage Logging: A private channel for private data

Forty wire telephone switchboard, 1907, Author unknown, Popular Science Monthly Vol 70, Wikimedia Commons.

A few months ago Crossref announced that we will be launching a new service for the community in 2016 that tracks activities around DOIs recording user content interactions. These “events” cover a broad spectrum of online activities including publication usage, links to datasets, social bookmarks, blog mentions, social shares, comments, recommendations, etc. The DOI Event Tracking (DET) service collects the data and make it available to all in an open clearinghouse so that data are open, comparable, audit-able, and portable. These data are all publicly available from external platform partners, and they meet the terms of distribution from each partner. Continue reading “Distributed Usage Logging: A private channel for private data”

DOI Event Tracker (DET): Pilot progresses and is poised for launch

Publishers, researchers, funders, institutions and technology providers are all interested in better understanding how scholarly research is used. Scholarly content has always been discussed by scholars outside the formal literature and by others beyond the academic community. We need a way to monitor and distribute this valuable information. Continue reading “DOI Event Tracker (DET): Pilot progresses and is poised for launch”

DataCite supporting content negotiation

In April CrossRef launched content negotiation support for its DOIs. At the time I cheekily called-out DataCite to start supporting content negotiation as well.

Edward Zukowski (DataCite’s resident propellor-head) took up the challenge with gusto and, as of September 22nd DataCite has also been supporting content negotiation for its DOIs. This means that one million more DOIs are now linked-data friendly. Congratulations to Ed and the rest of the team at DataCite.

We hope this is a trend. Back in June Knowledge Exchange organized a seminar on Persistent Object Identifiers. One of the outcomes of the meeting was “Den Haag Manifesto” a document outlining five relatively simple steps that different persistent identifier systems could take in order to increase interoperability. Most of these steps involved adopting linked data principles including support for content negotiation. We look forward to hearing about other persistent identifiers adopting these principles over the next year.

Having said that, this time I will refrain from calling-out anybody specifically…

Enhanced by Zemanta

DOIs and Linked Data: Some Concrete Proposals

Since last month’s threads (here, here, here and here) talking about the issues involved in making the DOI a first-class identifier for linked data applications, I’ve had the chance to actually sit down with some of the thread’s participants (Tony Hammond, Leigh Dodds, Norman Paskin) and we’ve been able sketch-out some possible scenarios for migrating the DOI into a linked data world.

I think that several of us were struck by how little actually needs to be done in order to fully address virtually all of the concerns that the linked data community has expressed about DOIs. Not only that- but in some of these scenarios we would put ourselves in a position to be able to semantically-enable over 40 million DOIs with what amounts to the flick of a switch.

Given the huge interest in linked data on the part of researchers and CrossRef members- it seems like it would be a fantastic boon to both the IDF (International DOI Foundation) and CrossRef if we were able to do something quickly here.

Anyway- The following are notes outlining several concrete proposals for addressing the limitations of DOIs as identifiers in linked data applications. They range in complexity/effort involved- with the simplest scenario providing minimal (yet functional) LD capabilities for just one RA’s members (CrossRef’s) and the most complex providing per-RA and per-RA-member configurability on how DOIs would behave for LD applications.

We’d appreciate comments, questions, suggestions, corrections, etc.

A: Simplest Scenario

What would need to be done?

  1. CrossRef implements a linked data service. For example, hosted at rdf.crossref.org.
  2. CrossRef recommends that any member publisher who wants to add rudimentary linked data capabilities to their site could simply insert some simple link elements into their landing Pages. So, for instance, for the article with the DOI 10.5555/1234567 in the Journal of Psychoceramics, the publisher would put the following in the landing page for the article:
<link rel=”primarytopic” href=”http://doi.crossref.org/10.5555/1234567″ /> 
    <link rel=”alternate” type=”application/rdf+xml” href=”http://rdf.crossref.org/metadata/10.5555/1234567.rdf” title=”RDF/XML version of this document”/> 
    <link rel=”alternate” type=”text/html” href=”http://www.journalofpsychoceramics.org/10.5555/1234567.html” title=”HTML version of this document”/> 
    <link rel=”alternate” type=”application/json” href=”http://rdf.crossref.org/metadata/10.5555/1234567.json” title=”RDF/JSON version of this document”/> 
    <link rel=”alternate” type=”text/turtle” href=”http://rdf.crossref.org/metadata/10.5555/1234567.ttl” title=”Turtle version of this document”/>

In the above snippet the HTML version of the document is the publisher’s existing landing page.

How it would work

  1. A sem-web-enabled browser would query dx.doi.org/10.5555/1234567 and get a normal 302 redirect to the publisher’s landing page. 
  2. The sem-web-enabled browser would sniff the page for the link elements and retrieve the representations it wanted from rdf.crossref.org
  3. The returned document would contain an appropriate representation of the metadata that the publisher has deposited with CrossRef. It would also assert that:

doi.crossref.org/10.5555/12334567 owl:sameAs dx.doi.org/10.5555/1234567 .
dx.doi.org/10.5555/12334567 owl:sameAs info:doi/10.5555/12334567 .

info:doi/10.5555/12334567 owl:sameAs doi:10.5555/1234567 .

Alternatively, the publisher could implement their own linked data support on their own domain using whatever appropriate method they want. So, for instance, a larger publisher could support content negotiation at their site and return different/enhanced metadata, etc.

Pros

  1. Doesn’t require changes at DOI/Handle levels
  2. Is easy for publisher to opt-in or opt-out
  3. Requires minimal development on the part of CrossRef.

Cons

  1. Only applies to CrossRef DOIs.
  2. It depends on publishers taking action. Might be a long time before publishers add the needed links to their landing pages or support content negotiation.
  3. DOI system is still not strictly LD compliant (e.g. it is returning 302 redirects. Naive sem-web browsers might ‘stop’ after getting a 302. Should ideally use 303s, content negotiation, etc.)
  4. Doesn’t work for DOIs that currently bypass landing pages and which go directly to content.

B: Simple + IDF Global Semantic Compliance

What would need to be done?

  1. Same as “Simplest Scenario”
  2. IDF globally changes dx.doi.org to return 303 redirect

How would it work?

Same as Simplest Scenario, except that, because sem-web-enabled browser had been told it was being redirected to a NIR (via the 303), it would presumably be more likely to continue.

Pros

  1. All DOIs conform to expectations for LD identifiers
  2. Easy for publisher to opt-in or opt-out
  3. Requires minimal development on part of CrossRef
  4. Requires minimal work (?) on part of IDF

Cons

  1. Requires global change on part of IDF. Global change might conflict with requirements of other RAs.
  2. It depends on publishers taking action. Might be a long time before publishers add needed links to their landing pages or support content negotiation.
  3. Doesn’t work for DOIs that currently bypass landing pages (e.g. OECD spreadhseets, UICR datasets, etc.)

C: Simple + IDF Global Semantic Compliance + RA CN Intercept

What would need to be done?

  1. Same as “B: Simple + IDF Global Semantic Compliance” Scenario
  2. IDF  changes dx.doi.org to redirect content-negotiated dx.doi.org queries to RA-controlled resolver depending on the preferences of the RA.
  3. RA implements DOI resolver (e.g. dx.crossref.org) that supports content negotiation. RA allows its members to specify to the RA  that they want either:
    1. RA to forward all requests to the member’s site.
    2. RA to “intercept” content-negotiations for non-HTML representations and direct them appropriately (e.g. return appropriate representation from rdf.crossref.org)

How would it work?


Pros

  1. All DOIs conform to expectations for LD identifiers
  2. Allows RA to potentially LD-enable its members very quickly.
  3. Easy for ra-members to opt-in or opt-out
  4. Requires minimal development on part of CrossRef
  5. Would even work for DOIs that bypass landing pages

Cons

  1. Requires global change on part of IDF. Global change might conflict with requirements of other RAs.
  2. Requires change to add decision logic implementation on part of IDF. 
  3. Requires development of RA resolvers that implement per-member resolution logic (note- this would probably actually be done at DOI level)

D: Simple + IDF Selective Semantic Compliance + RA CN Intercept

What would need to be done?

  1. Same as Simplest Scenario
  2. IDF  changes dx.doi.org to return either 302 or 303 redirect depending on the preferences of the RA.
  3. IDF  changes dx.doi.org to redirect content-negotiated dx.doi.org queries to RA-controlled resolver depending on the preferences of the RA.
  4. RA implements DOI resolver (e.g. dx.crossref.org) that supports content negotiation. RA allows its members to specify to the RA  that they want either:
    1. RA to forward all requests to the member’s site.
    2. RA to “intercept” content-negotiations for non-HTML representations and direct them appropriately (e.g. return appropriate representation from rdf.crossref.org)

How would it work?

Pros

  1. Allows RA to potentially LD-enable its members very quickly.
  2. Easy for ra-members to opt-in or opt-out
  3. Requires minimal development on part of CrossRef
  4. Would even work for DOIs that bypass landing pages

Cons

  1. Only some DOIs conform to expectations for LD identifiers
  2. Requires change to add decision logic implementation on part of IDF. 
  3. Requires development of RA resolvers that implement per-member resolution logic (note- this would probably actually be done at DOI level)