Showing posts with label DOI. Show all posts
Showing posts with label DOI. Show all posts

Wednesday, June 14, 2017

Digital Object Identifiers for StatCan data

Background
Back in 2011 and then again in 2015 there were some questions on the list about DOIs and StatCan data.  At one point it was asked specifically if STC had considered registering DOIs (it was right after DataCite Canada had launched and was being tested) and later [after the Abacus group (SFU, UBC, UNBC, and UVic) made the transition to Dataverse, it was] asked about assigning DOIs and [a] discussion ensued. Now UNB is also going to use Dataverse to offer up research data and were thinking to use it for secondary data as well, but  we have discovered that we have no choice about DOIs. That is, Dataverse is Harvard's resource and they officially indicate that DOIs are assigned. Period. Since Dataverse assigns DOIs automatically, we want to register them (or what's the point?) and are trying to figure out if there is a way that we might be able to force those DOIs to match STC registered DOIs if there were such a thing.

Question 1
If that's not possible, then, the question becomes: what is StatCan's official line on other institutions effectively assigning DOIs to their (STC) data?

Question 2
So to STC employees I would ask a) is Statistics Canada registering DOIs or do they plan to and b) is Statistics Canada concerned about multiple universities (or other organizations) assigning DOIs?  Perhaps we've missed something in the Dataverse documentation, but it really doesn't seem like we have a choice about the assignment of DOIs.

Question 3
My question to the DLI community is how do you deal with this issue (i.e., that Dataverse doesn't give you the choice of whether it assigns a DOI and it doesn't look like we can suppress the info) at your institution?  

Answer 1
  • At UBC, we mint DOIs only for research data, not licensed datasets. 
  • We do not mint DOIs via Abacus Dataverse but via our discovery layer - Open Collections - https://open.library.ubc.ca/, which allows us great flexibility for DOIs minting
  • The newest Dataverse version - 4.6.2 allows to mint handles in addition to DOIs, which might solve the UNB problem. We have collaborated with Harvard to offer that...
    • i.e.: https://dataverse.org/blog/dataverse-462, very new , just released last week or so...So good timing. I was working with Harvard on that for more than a year. Developed by DANS (our Dutch colleagues).More on Github - https://github.com/IQSS/dataverse/milestone/61?closed=1
  • We had to develop our entire DOIs GUI and a pipeline as Datacite Canada was not flexible enough for us. Here is more information - http://researchdata.library.ubc.ca/plan/get-dois/
  • By now we have minted more than 215,000 DOIs for our digital assets (out of 274K in Canada - https://stats.datacite.org/?fq=allocator_facet%3A%22CISTI+-+National+Research+Council+Canada%22&#tab-datacentres)
  • We have assisted multiple schools in their DOIs work, namely uOttawa, BC ELN, Guelph, McMaster, VIU and many more...

Answer 2
We [StatCan] have reached out internally to obtain more information.  Statistics Canada is collaborating with NRC’s DataCite to register DOIs for its aggregate data on the website.  While this will still take some time, progress is underway. I brought forth the concern from the community registering their own DOIs for statcan data. Statcan consulted with DataCite representative that indicated that the current best practice thinking is that multiple DOIs are accepted, as long as they are from different clients. There are good use cases for both registrations, so that another DataCite client with a different prefix will be able to assign a DOI to a copy of Statcan content stored in their repository.  

Informally, I was informed it would be ideal if repositories linked to the official Statcan DOI once available.  As the authors of the data, this would ensure that users are directed to the current and authoritative source.  At some point in the future, they have agreed it would be beneficial to have a conversation with stakeholders of the community.  We, the DLI, are continuing to have conversations with internal stakeholder regarding the potential for registration of items, such as PUMFs.
[StatCan] is interested to learn more about the communities perspectives, please share your comments on the list.


Further comments
Hi folks, I’m attending the Dataverse Community Meeting this week, the release which was discussed below allows support for Handles OR DOIs as the persistent identifier for your Dataverse instance, not both (if I’m clear on this, haven’t actually tested it yet). In our case, I believe we would want to support the option for selecting either a handle or doi on a dataset by dataset basis in the same DV instance  – see this use case ticket here: https://github.com/IQSS/dataverse/issues/3623

I’m also interested in coming up with a coordinated solution for registering DOIs for STC data, including aggregate data available to us via the DLI. I’m happy that STC and DLI may be assigning DOIs in the near future, which is a step forward. We’ve discussed this at length within the OCUL community and I hope it can be a topic of discussion at our national training next year in Montreal (I’m assuming this is happening). 

Here are some options we’ve explored/discussed:
  • publishing STC data w/ DOIs (argument that these are our access copies; according the DataCite BP)
  • publishing STC data w/ DOIs but unregistering these DOIs using the DataCite API
  • publishing STC data w/ Handles (however, not technically possible in Dataverse yet)
  • publishing STC data w/ internal Dataverse identifier (same as above)

I have some questions for other folks in Canada that might be helpful for our conversations…If we were to coordinate loading of STC data w/other Canadian universities 
  • what data are you loading?
  • can we harvest one repository? 
  • can we harvest DLI’s repository? 
For now our PUMFs will remain in our Nesstar repository, but we are in the process of releasing all non-PUMFs in our SP Dataverse. 

Further comments
The issue of duplicate DOIs is becoming a concern not only in our environment but elsewhere as well. At least, STC is now considering assigning them to aggregate data – a first step. It was also good to hear what is [happening] out west as this is an issue we have to consider for ODESI and the Scholars Portal Dataverse. There should be some interesting information coming out of  the DataVerse conference – I look forward to hearing about it from anyone else who is attending.

Tuesday, March 24, 2015

DOIs for Stats Canada and Co. Datasets

Question

While trying to work on solutions for research data identifiers (DOIs in our case), we have stumbled upon some unique issues with the Stats Canada datasets you are all familiar with.

In a nutshell, we have more than 27,000 data files that our ever amazing Paul Lesack is managing for four BC schools for licensed data in our Dataverse <http://dvn.library.ubc.ca/ dvn/>

As we are planning to assign DOIs to any Dataverse files on our end...we are wondering about the federal government datasets. These do not seem to have any unique URLs (e.g. handles or DOIs). If we assign DOIs to these files on our end and say five other schools assign DOIs to the same datasets, we will have a small army of DOIs floating around for the same data.

Working with NRC/CISTI as a Canadian DataCite agency, we seem to be the first ones to deal with this issue. Mind you, we also plan to assign DOIs to all our digital objects (CONTENTdm, DSpace, etc). Still, we are really interested to hear your thoughts.

I have recommended to CISTI to proactively approach Stats Can and Health Can (and others) about assigning DOIs to content they produce. Would you agree to this practice?

And please excuse me for my ignorance as I am not really a data librarian but trying to build a research data service for our large campus.

Responses

I would agree with you that CISTI should proactively approach STC and other gov. depts. about DOIs. It is an issue which was first mentioned a couple of years ago to DLI and which needs to be addressed.

You have hit the nail on the head when you state that different schools/repositories cannot all be assigning unique DOIs to the same STC dataset. At Scholars Portal we have discussed this before with respect to <odesi> and we, like you, had concerns with the idea of there being many DOIs floating around for the same data.

Personally, I have no problem with assigning DOIs to the locally hosted datasets, but then again these datasets will have numerous persistent identifiers depending on the hosting institution. I 
am guessing the DLI Nesstar server would not work at this stage for this? 

However, in regards to how DOIs work, it is my understanding that a DOI registered by Statistics Canada would resolve to a webpage maintained by them. So, any time this DOI were cited somewhere it would point back to Statistics Canada, which makes sense. Except what if you specifically wanted to point people back to the copy of the dataset housed in your Dataverse? Would you ever want to do that? For instance, we recently noticed that at least one third party (in the US) has registered at least one DOI for StatCan data: <http://data.datacite.org/10.6068/DP14A4B06A47153>. Not sure if this is just poor practice, or if it is something that will occur on a regular basis.

Dataverse allows us to host a variety of licensed and research data. Even data on social housing of dairy calves ​<hdl.handle.net/11272/10178>. This is the blurb on NRC's website re DataCite: "NRC is a founding member of DataCite and is its DOI allocation agent for Canada." I believe this is the direction many are going with respect to DOIs and assignment for multiple iterations/versions/copies of data. However, there needs to be some authority record or reference identifier to properly identify data sets and the study.

At last year’s IASSIST conference there was a panel discussion on data discovery that mentioned the use of data set identifiers and I believe this issue was discussed by the panelists <http://www.library.yorku.ca/cms/iassist/program/sb4/#sb4o>. Many data organizations host the same data sets (in terms of data ‘s content), but the iterations/copies to go for access are different (UBC hosted, vs. <odesi>, vs. DLI etc.). There is no metadata clearinghouse for data sets, DataCite being the closest thing.

I think there could be a practice in place where the original data producer assigned the DOI for the data set, and that DOI reference was carried forward into other iterations of the data set for reuse, with some customization of the identifier for different access points. Likewise, I believe there is a way to customize the DOI to include reference to other identifiers such as STC catalogue #s or the IMDB, etc., in cases where there is no DOI. For example the ISBN > DOI integration <http://www.doi.org/factsheets/ISBN-A.html>.

However, we need to be very careful with regard to assigning DOI especially where we may be assigning multiple DOI identifiers to the same object. Such a practice is discouraged by the doi.org.  The doi Handbook <http://www.doi.org/hb.htmlhttp://www.doi.org/hb.html> states:

"Each DOI® name is a unique "number", assigned to identify only one entity. Although the DOI system will assure that the same DOI name is not issued twice, it is a primary responsibility of the Registrant (the company or individual assigning the DOI name) and its Registration Agency to identify uniquely each object within a DOI name prefix.

Uniqueness (specification by a DOI name of one and only one referent) is enforced by the DOI system. It is desirable that two DOI names should not be assigned to the same thing."

Likewise, in regards to republished or duplicate datasets:

"We strongly recommend that DOIs be created only for ‘original’ datasets, not duplicate datasets. There may at times be a need to deposit duplicate copies of a dataset in multiple data centres, for example where a project has been funded by multiple funders and each funder requires deposition in a different data centre. If possible we would suggest identifying the primary version of the dataset and assigning a DOI to this version only. Where there is an unavoidable need to publish a dataset in different locations each with a separate DOI, the metadata for each appearance of the dataset should indicate the association."

<http://cisti-icist.nrc-cnrc.gc.ca/obj/cisti-icist/doc/datacite/datasets.pdf> -- scroll to the bottom..