ECONOMIC INQUIRY | Data Availability Policy
Economic Inquiry requires that authors of all published empirical papers submitted to the journal after December 1, 2021 provide enough detail on their work so that their results can be replicated. It is expected that, prior to publication, not upon initial submission, data from empirical papers will be uploaded into a suitable archive and that instructions be included with the data necessary for an interested party to reconstruct key tables and figures from the published manuscript. For any paper that involves gathering a unique data set, this includes economic experiments, surveys and similar data gathering exercises, all instruments used in gathering that data should also be provided in that archive. The replication materials will be linked to the published paper.
The WEAI has a sponsored data archive at ICPSR in which all authors can upload their replication materials free of charge after the paper has passed the review process (visit the WEAI repository for more information and instructions). Authors should expect to use WEAI's archive site but can, if necessary, use an alternative archive site. Any alternative site must be approved by the editor to ensure that the alternative archive site meets the journal’s requirements for data integrity.
After a paper receives a conditional acceptance, the authors will be asked to upload their data archive and provide the study number or URL of its location. A member of the EI Data Team will review the data archive to ensure that it meets the journal requirements. Authors are encouraged to pre-review their own data package using the same checklist members of our Data Team will use to verify that your submission meets our requirements. Prior to uploading your data archive, you can go through the same checks our team will to make certain that someone reviewing your archive would be able to check “Yes” to the relevant questions. If there are any questions where “No” would be selected, you can then make certain that your README file explains the reasons for that omission. By pre-reviewing your own data archive to ensure compliance, you can lessen the time spent in the data review process and get your paper published more quickly.
Once your materials are submitted, our Data Team will review them. If the materials provided are not complete, authors will receive communication from the Data Team pointing out what element of the checklist your archive was incomplete on with indications for how to rectify the issue. This will continue until an acceptable archive is produced. Under normal circumstances, this process should not take long to complete. If authors fail to comply with the requirement for providing replication materials within six months or cannot provide an acceptable archive within three submission opportunities, the journal maintains the right to revoke the conditional acceptance and reject the paper. As with our editorial decisions, we also seek to limit the number of rounds of submissions of data packages and so authors should pre-check their data packages carefully prior to submission to make sure that they should meet the requirements so that multiple submission rounds are unnecessary. Should replication materials turn out to be problematic after publication has occurred, the editor will deal with cases as they arise and in serious cases papers may be retracted by the journal.
Replication packages must include the following elements:
- READ ME. A summary file (preferably in pdf) describing the contents of the replication package. It should explain the role and function of each file included and detail all software necessary to run the code as well as any additional add-on packages required. A simple explanation should be included providing instructions for how someone should run the code to generate the results, as well as an explanation for where the results can be found once the code is finished. We require that this file should follow the format as specified in the README Template provided by the Social Science Data Editors.
- Data. We expect all raw data to be included in the archive, and typically all the processed data should be included as well. Submitting only processed data is not sufficient. Processed data files are those that have been modified from the source in any way, including files that have been lightly cleaned, reformatted, transposed, filtered, or merged with other files. Additionally, all sources must be fully described with sufficient detail so another researcher can directly obtain the data from the source. Authors are responsible for ensuring that they have the legal right to use and redistribute any data they deposit in the archive. Data that cannot be shared is discussed later.
- Code. We also expect all code to be included in the archive, including the code that creates processed datasets, as well as code to generate results in the manuscript and appendices. When some final or intermediate datasets are created manually, the procedure to construct them from the raw data must be fully described so that any interested party can follow the instructions and recreate that processed data and/or result.
- Original Data. For any projects that involve the creation of an original data set via surveys, experiments, or similar methods, authors should provide full details on the methods used in that process. This involves providing data-gathering instruments, experiment programs, instruction scripts, and so on, including a description of how these materials were used in gathering data.
- Web-scraped data. WEAI and Economic Inquiry are bound by US law, and we take cues from the legality of using scraped data from the AEA Data Legality page. As noted there, web scraping may breach terms of use of a website, but its legality is not settled as of January 2023. So unless the law changes, we will consider the data to be legally obtained and require that authors deposit both the raw data as well as code to collect it. Thus, for data that has been scraped using computer code, we expect a copy of the raw data (e.g., HTML, CSV, etc.), and code to scrape the data to be included in the archive. We will not test this code (and if it is called up somewhere from other code, the call should be commented out but highlighted in the README) as the website may not even have the same information/structure. However, any further code that generates processed data or directly any results from this data is expected to be functional.
- Simulations. For any projects that involve simulations or computational elements, the code generating those calculations should be included, as well as an explanation of how one would run the code. If the simulations require parameter values from the literature, a table should be included listing those parameters, their values, and a citation for the parameter value. If the parameters are based on moments calculated using other data, that data and code to compute the moments fall under earlier instructions about data and code.
- Restricted data. It is expected that some data sets may be proprietary, confidential, or otherwise restricted for redistribution and hence cannot be included in the archive. If that’s the case, the precise reason for the data being restricted should still be explained in the cover letter when submitting. Additionally, the following will then apply.
- Even if the data is confidential, full details about how another researcher can obtain this data must be included in the README. The information should include contact information (website, email address, postal address, office, etc.), what to request (name of series, coverage, etc.), and the names of exact data files initially obtained. Where applicable, a sample letter of request for data should also be included in the archive, for instance, if the data was obtained via Freedom of Information.
- In cases where data is commercially available to anyone but obtained via a restricted database or website or other similar special data availability arrangements of the authors' institutional arrangement, again those should be fully described in the README along with contact information, and what steps to follow to extract the same series and coverage.
- If there is some other special arrangement with a party to obtain or use their data and it is not available to others, a Data Use Agreement (DUA) or other similar documentation should be shared with the Data Editor for verification. This does not need to be included in the archive.
- In all such cases, authors will be asked to provide synthetic/fake data. See next point.
- Synthetic data. If data is restricted and cannot be shared, it is expected that authors will provide synthetic/fake data in its place. The primary purpose is to test and understand the code that is shared. The data need not preserve moments of the true data to mimic the results of the paper (they can if the author prefers), but all the same variables, number of observations, and data types should be preserved to match the true data so the code works on it. In some cases, it may be that raw data is restricted due to confidentiality or due to the proprietary nature of the data, but the analysis is on some higher aggregated version where those concerns don’t apply. If that is the case, the authors should consider providing raw synthetic data and true processed data.
- Personally Identifiable Information (PII). Authors are responsible for removing all PII where appropriate as well as following the rules established around data sharing by their institutions' ethics board. The Data Team does not verify that deposited files are free of PII; responsibility for identifying and removing any PII that should not be shared rests with the authors. However, if we find any cases where such information is included in the data files and should not have been, we will notify the authors. If it requires unpublishing the original published archive because data included PII that was not supposed to be there, authors are expected to provide an alternative anonymized version. Failure to do so can result in retraction of the paper.
- Legality of data. As a general rule, it is automatically assumed that all data used in the paper and provided in the archive are legally acquired, and this applies even if no specific declaration to that effect is made in the README or elsewhere. Ex-post discovery of illegally acquired data can lead to retraction of the paper and unpublishing of the data archive. There could be exceptions to this where even the author did not know that it was illegally acquired and they obtained it in good faith from a third party. Such cases will be handled on a case-by-case basis. Another important (and rare) case could be where the authors knowingly use illegally acquired data but make a compelling case for using that data in the public interest. If that is the case, this should be disclosed upfront at submission time, and discussed with the Editor-in-Chief directly, who may then give an exemption if it is legally possible and if there is a way of establishing the authenticity of the data, which may require further checks and/or discussion with the Data Editor.
If there are reasons that you would be unable to comply with this policy, those reasons should be explained when the paper is submitted. Waivers to this policy can be allowed when appropriate at the discretion of the editor.