email us at Bioinformaticsolutions @ gmail . com

Showing posts with label Clinical Data. Show all posts
Showing posts with label Clinical Data. Show all posts

Friday, February 8, 2013

Recent Examples of Infrastructure (Databases and Analytics) Engagements in Healthcare (IBM, Truven, Oracle, ATG)

For those of you who enjoyed our reports on informatics platform development in the UK (use the search feature on the blog to look for posts referencing 'NOCRI'), we're going to be focusing again on enterprise platforms which support service delivery and research with a global lens in the coming months. To whet your appetite for upcoming posts here's a selection of projects we've come across in the last week:
 From HealthDataManagement.com: Truven Health Analytics has introduced a new suite of products for statewide health information exchanges, called HIE Advantage Analytics
The suite is designed to enable public health officials to access and analyze operational and clinical statistical data in a state’s HIE. The West Virginia Health Information Network is an early adopter.
HIE Advantage uses near real-time clinical data from providers, claims data from the Centers for Medicare and Medicaid Services, and data from other sources that are stored in state HIE repositories. The goal is to monitor community health status while improving outcomes, according to Truven Health, previously the health care business of Thomson Reuters.
Reports analyze prevalence, process of care and outcome metrics for specific diseases, as well as rates of screening and preventive care to identify communities at higher risk for poor health status. More information is available here.
From Investors.com: The University of Texas MD Anderson Cancer Center Selects Oracle Applications and Technology as Part of Platform to Help Transform Cancer Care
·         MD Anderson, one of the world's most respected centers devoted exclusively to cancer patient care, research, education and prevention, has selected Oracle Health Sciences applications and Oracle technology as the foundation for an organization-wide analytics initiative designed to enable a new generation of personalized cancer treatment that improves outcomes. The platform will also support the center's renowned Moon Shots Program, an unprecedented effort to dramatically accelerate the pace of converting scientific discoveries into clinical advances that reduce cancer deaths.
·         Oracle applications and technology will power the enterprise analytics initiative, one of the program's platforms. The new platform will bring together clinical, genomic, financial, administrative and operational information from internal and external sources to yield insights that drive care innovation and optimize operational efficiency.
·         To achieve its goals, the center, ranked first for cancer care in U.S. News & World Report's "Best Hospitals" survey for seven of the last nine years, will deploy a wide range of Oracle solutions, including Oracle Enterprise Healthcare Analytics and Oracle Translational Research Center
·         MD Anderson, which sees data growth of 30 percent to 40 percent annually, is also deploying Oracle Database and Oracle Business Intelligence Enterprise Edition.

From Executive Biz IBM to Provide Int’l Medical Research Facility with IT Infrastructure, Application Support
IBM and the Dmitry Rogachev Clinical Center have entered into a contract of agreement to deploy PureFlex integrated systems in the Clinic’s facilities to bolster IT support, according to an IBM statement.
“By 2015, we expect to increase the volume of clinical tests five times and to be able to cover 5,000 primary patients generating over a petabyte of medical data,” said Igor Pyatnitsa, head of operations at the Russian Federal Scientific-Clinical Center of Pediatric Hematology, Oncology and Immunology.
“This is valuable data which we must effectively store and manage for medical and legal reasons. IBM’s PureFlex systems help us to do this effectively while controlling costs and ensuring the highest levels of data security,” he added.
The agreement is an extension of an existing contract for the IBM PureSystems roll out.
The center focuses on finding treatment for blood disorders, cancer, immune system diseases, and other diseases. It is a part of the Russian Ministry of Health and a chief collaborator on more than 400 projects, 100 clinical trials and 20,000 medical tests per year.
This will be the first time PureFlex will be used in Russia. PureFlex technology will assist the center with installing fast systems and critical medical applications.
It will also help in organizing their extensive medical data repository through an automatic locator using existing applications. The hospital hopes that this will lead to better collaboration and field research.
Andrey Filatov, Director, IBM Systems and Technology Group for IBM in Russia and the CIS said the PureFlex Systems is meant to provide a platform for Russian healthcare and medical research development.
The offering is tuned for cloud computing and can consolidate more than 100 databases on a single system and helps to rapidly deploy medical applications.

From Executive Biz: Allied Technology Designing, Analyzing Army Medical Research Projects [Databases, Analytics]
Maryland-based information technology and engineering provider Allied Technology Group has won a five-year contract to help a U.S. Army medical research facility in Silver Spring design, analyze and report on projects.
ATG says its statisticians and public health analysts work with the Walter Reed Army Institute of Research ("the largest and most diverse biomedical research laboratory in the Department of Defense") to include and exclude criteria, select study subjects, develop analytic databases and analyze statistics.
The company will provide the institute a team of epidemiologists, biostatisticians and administrative personnel for analysis, collecting and entering data, managing databases and programming computers.
Since 2007, ATG says staff members have authored and co-authored articles for 30 peer-reviewed publications and 34 scientific presentations on their work at the institute.

    Monday, December 3, 2012

    European data assets catalogue & eHealth Infrastructure - analyses of progress

    We'll be looking at access to European patient-level data for research purposes over the next few months and have hit on a couple of online resources we thought worth sharing with you.

    The first is a listing of data assets by country (and some international assets) with clear concise descriptions of their content.This is a great starting place for anyone looking to understand the large-scale datasets which might support various kinds of clinical, epidemiological and commercial research. The below is an excerpt, the entry for the Netherlands, to whet your appetite:

    Netherlands
    Health records
    • Integrated Primary Care Information (IPCI) information from electronic patient records of 150 GPs covering more than 1 000 000 patients
    • Pharmo Independent research organization for drug use and outcomes (including cardiovascular, metabolic disease, oncology and autoimmune disease, respiratory disease, and mother and child health). Overall it covers 2 million residents in the Netherlands and around 200 000 patients linked to GP patient records. Other data includes:
      • Community Pharmacy database (CPD)
      • Clinical Laboratory File (CLF)
      • General Practitioner database (GPD)
      • Dutch Pathology Registers (PALGA)
      • Hospital Pharmacy database (HPD)
      • Dutch mortality statistics (CBG)
      • National Dutch Hospital Registration (LMR)
      • Perinatal Registry (PRN)
      • Eindhoven Cancer registration (IKZ)
    Health statistics
    • GIP database from Health Care Insurance Board (Free online) covering outpatient drug utlization for 85% of population
    The second resource is a series of reports by country, together with a summary across the continent, of eHealth Infrastructure initiatives in Europe. The final report is dated January 2011 but the speed with which these state and EU funded initiatives are able to progress is such that we suspect these analyses still stand and can be considered fairly current. The summary report is a good document to start with to understand the scope and purpose of the analysis but the individual country briefs are invaluable in their discussion of, for example, the data integration issues thrown up by Spain's regional administration, or the legal impediments to cross-border data sharing.

    Thursday, June 21, 2012

    Google BigQuery - Hospital Episode Statistics data analysis made easy (ish)

    We had a visit from PA consulting last week, who were very excited about their application of Google's new Big Data solution, BigQuery, to the vast dataset which is the UK's Hospital Episode Statistics (HES).

    We've had to wrestle with HES data before and come off the worse for it - and that was looking only at one year's Inpatient data - about 18 million lines of data. We encountered all kinds of issues in the setup of our bespoke database, created to hold and analyse this data, starting with the published data dictionary's divergence from the fields in our extract, taking in the discovery of duplicated unique episode identifiers (see our post on the response from the NHS Information Centre) and ending with our discovery that the processing power required to run some of our composite queries in a timely fashion was beyond our meagre infrastructure (this was at the now defunct National Cancer Research Institute's Informatics Initiative).

    We could really have used the facility which BigQuery is set to provide. The guys from PA described how they had obtained the entire start-to-finish HES dataset across all three areas of collection (inpatient, outpatient and A&E) and loaded this into BigQuery (this being the most arduous part of the process, the data arriving on 27 DVDs and taking a couple of weeks to upload) prior to demonstrating the speed with which it was able to provide answers and how the data could be linked to google maps and google docs' spreadsheet application to dynamically produce visual and graphical analyses.

    BigQuery dynamically calls in servers to assist in the running of a query based on the processing power required and then releases them once the query has been executed. The result of being able to access Google's immense army of servers is that without any of the usual time-consuming optimisation (indexing etc.) which supports enhanced performance on traditional database technologies, the user can execute a query against billions of data points in seconds. If you're working to some degree 'in the dark', uncertain of how you wish to structure your data and what analyses you will require it to support, you can experiment on vast datasets without waiting hours for queries to produce results (or fail!) - a facility which would have delivered huge time-savings to us in our HES analysis.

    PA have a video describing their work and approach here and Google provide further information about BigQuery here.

    Monday, June 11, 2012

    Clinical data - toward a single dataset supporting Research, Service Delivery and Performance Management

    We were lucky enough last month to discuss with the Finance Director of a major London healthcare service, the use of enterprise data, from clinical to financial, to support performance management - and were struck by the extent to which the processes which support performance management analytics mirror those which support clinical research.



    We were put in mind of our conversations earlier this year with Oracle and other vendors whose healthcare data warehouse platforms and attendant applications were being demonstrated as solutions to both research and performance measurement problems - this in turn called to mind the comments we had heard from data experts coming into healthcare from other industries who could not understand why information of all kinds pertinent to the administration of healthcare, from genomic analyses to staff costs, were not seen as belonging to a single and vitally important information asset.



    The concerns of this London hospital in getting a handle on their data were strikingly similar to those of the research institutes we have spoken with - and indeed what they describe as performance management is really research by another name - when they correlate outcomes with treatment modalities and different packages of care, they are using their Business Intelligence architecture in many ways like a clinical research engine but with the addition of financial data .




    Their core issue is data quality - they currently have four main clinical IT systems - one for each borough subsumed into their organisation. Trying to use these to derive information even at the level of 'number of patient encounters' has not been straightforward. Although they didn't go into it, we should imagine that supporting those clinical systems must be an array of systems capturing e.g. pathology, radiology, cytology data at a more granular level.



    In addition, their financial data resides in three ledgers each with different coding - a 'consolidation nightmare' was how they described the move to a single ledger; painful but essential as they attempt to get a grip on expenditure.



    The introduction of Service Line Reporting of income and expenditure (in order to assess profit by service) is driving their data validation and data quality improvement - but their clinical operations requirement to deliver an integrated service, developing 'packages of care' rather than looking at individual activities related to the same condition in isolation, is also dependent on quality and timely data - both financial and clinical performance management require the facility to benchmark accurately.



    Touching on other IT issues their organisation faces, they mentioned that the mobile workforce are not well supported by technology - "it's still a surprise that no-one has developed a good mobile working solution for healthcare in the UK". Their words, not ours! They did suggest complications which we hadn't thought of hitherto, coming from a tertiary care background as we do, for example nurses on home visits may not be able to work online, thus need data stored locally to upload later which means on-device storage of personal identifiable data. We can't believe that is still an issue from a technological perspective, however, we can imagine that risk-aversion in respect of personal data in the healthcare industry is dampening demand for solutions - anyone care to offer a more informed opinion?All comments welcome!



    We were really interested to hear how they addressed the development of their Business Intelligence capabilities - developing KPIs to meet significant and varied requirements from different commissioners. Their previous dashboards had been provided by NHS London - their goal at that time had been to ensure compliance with standards and regulations but as they have matured they now have to develop their own dashboards to meet more complex internally-driven reporting requirements. To do this, they have their own Performance and Information team who work on data collation and aggregation - creating Performance Packs - which provide detail by Service line under headings such as Operations, Quality, Finance, Workforce etc. supported by detailed analysis across the board from hard to soft data.



    The capacity to present a performance summary across directorates has led to internal competition which is already leading to performance improvements. Their next steps?



    the P&I teams are looking to automate the production of their performance packs and to create an overall dashboard for the organisation, leading in turn to a Balanced Scorecard.



    Having realised that their previous KPIs and the systems which provided the data were inadequate, they embarked on a redesign of their processes by, and in this order!:



    • Defining the goals / purpose / vision of the organisation

    • Asking what information they need to support the delivery of these

    • Asking what KPIs would adequately measure their delivery

    • Then developing the systems which support the provision of the answers to the above

    They are no longer looking simply at meeting regulatory reporting requirements - but at using their data internally to drive their performance - they are setting up data quality fora - having an external data quality audit and linking their output data to their income - and beginning to realise the benefits of placing data at the heart of the organisation.



    Monday, May 14, 2012

    Clinical Practice Research Datalink courts potential users

    We received on Friday our invitation to the Clinical Practice Research Datalink (CPRD) Users Meeting that will be held at the MHRA offices in London on 24th May 2012, consisting of a series of short presentations by representatives of CPRD and external partner organisations to give "further insight, and an opportunity for input, into the current and future aims of CPRD."

    The background we have covered severally before (see here and here): "CPRD is the new English NHS observational data and interventional research service, jointly funded by the NHS National Institute for Health Research (NIHR) and the Medicines and Healthcare products Regulatory Agency (MHRA). It combines the piloting work of the Research Capabilities Programme (RCP) and the existing General Practice Research Database (GPRD)."

    CPRD services are designed to maximise the way anonymised NHS clinical data can be linked to enable many types of observational research and deliver research outputs that are beneficial to improving and safeguarding public health. CPRD will act to provide services to a wide range of researchers and the aim of the Users Meeting will be to ensure that its plans meet the needs of the broad cross section of researchers in academia, the NHS and commercial companies both in the UK and globally.

    The specific topics that will be covered on the day include:
    • Pragmatic and Phase III - IV clinical trials
    • Multidimensional data quality
    • Hospital prescribing data
    • Models for linkage
    • Disease and patient group data marts
    The day will provide "several opportunities for potential users to raise and discuss their priorities and requirements" to ensure CPRD meets researchers' needs.
    It is possible that we won't be able to attend ourself so if any of our readers are planning on going along, let us know and we'll get in touch to see if you want to post some feedback on these pages!

    Monday, May 7, 2012

    Managing Research Data - the Joint Information Systems Committee Programme

    In covering the BRISSkit vision for a cloud-based open-source "research application as a service" at the end of last year, we mentioned the Joint Information Systems Committee's (JISC)  Managing Research Data workstream about which we've been hearing a lot more recently and thought we should pass on the basics and a link for your own browsing:

    JISC are a non departmental public body who support higher education and research in the UK by providing advice on the use of ICT - their Managing Research Data workstream supports both good data management and the sharing of data "for the benefit of UK Higher Education and Research". Their work in this area focuses on infrastructure, practice and skills:

    • piloting essential research data management infrastructures within institutions and for distributed research groups
    • improving practice in research data management planning
    • developing tools to help institutions plan their research data management practice
    • encouraging the publication of research data and demonstrating the benefits of improved methods for citing, linking and integrating research data
    • and, stimulating the acquisition of appropriate skills, among academics and research support staff in Universities
    Follow this link to their page which contains further information about their internationally recognised Digital Curation Centre and more information on the five strands of the programme which include projects, planning, tools and training.

    Wednesday, April 4, 2012

    Clinical Practice Research Datalink is finally here - or is it?

    The new Clinical Practice Research Datalink about which we have blogged much in the past has finally arrived (http://www.cprd.com/intro.asp) amid a certain amount of fanfare - see this from Pharma Times, this from PMLive and this from GP magazine.

     You may notice that the GPRD pages now redirect to this site and to a certain extent, this is largely a rebranding exercise at the moment. Behind the scenes a team at the DoH are trying to ensure that major data sources are willing and able to engage with this initiative but the speed at which they come online remains to be seen. We'll fill you in further on plans for a researchers data-catalogue interface as this project advances - and if CPRD are not offering that just yet, perhaps the MRC are - we'll get you up to date with the MRC's Data Support Service before the week is out.



    Wednesday, March 21, 2012

    Clinical Practice Research Datalink edges nearer

    We were interested to find today this URL for the new Clinical Practice Research Datalink about which we have blogged much in the past. Click on it and you will see the screen below (click the picture to zoom in) and thus be able to sign up for news of it's development. April seems pretty close now though we were aware that the Department of Health had set out some pretty aggressive deadlines - we hope to be surprised (pleasantly) come April Fools [note also that the MHRA appear to be hiring now for data specialists...]

    Tuesday, March 6, 2012

    Wales - what's their secret? Delivering successful healthcare informatics - care records

    We blogged in the past about the fact that Australia, Scotland and Wales had stolen a march on the English in developing data linkage facilities for healthcare research and had heard this blamed on the heterogeneity of our healthcare systems. We now read in eHealth Insider that "Wales is also having success with sharing patient data via its Individual Health Record. More than 300 GP practices have switched on access to the record, which is available to emergency care providers. It includes demographic information, medication, allergies, test and x-ray results, and medical problems from the past two years."


    A great comment on the eHealth Insider article compares the modest investment behind the Welsh achievement to that behind the English National Programme for IT with its "modest return".

    Wednesday, February 22, 2012

    General Practice data for research

    We have just read, courtesy of eHealth Insider that "the NHS Information Centre is on the verge of having all GP clinical systems suppliers signed up to the General Practice Extraction Service. Emis, Microtest, iSoft and INPS have signed up to extract and communicate data to the NHS IC, and TPP expects to have a contract signed within weeks."


    For more information on the General Practice Extraction Service (GPES), see the NHS IC GPES page which currently only mentions the contract with EMIS and whilst it describes the value of GP data, does not allude to its use for research but only in commissioning despite asserting that "GP patient records are the most complete record of a patient's health within the NHS. They comprise a wealth of information about patient care, the prevalence of diseases and treatments given."

    So where does this leave services like GPRD, THIN and QRESEARCH. Well, we know that GPES won't be fully up and running for another year at least, and that GPRD will become the CPRD  (Clinical Practice Research Datalink) if all goes to plan. In the interim, as per this from the Health Protection Agency, GPRD and QRESEARCH will still be providing primary care data for research, but what will the landscape look like by the end of 2013?

    Tuesday, February 21, 2012

    Patients' views on participating in medical research - Part 2: Engagement and Participation

    Our last post here began our summary of recent work on patients' attitudes to participating in medical research - part of our primer on the current state of play in clinical research in the UK selecting choice elements from recent reports; check out the footnotes for interesting sources to follow up. We include in this a look at the influence of media coverage of issues pertaining to data security on patients' attitudes and behaviour.

    As always, let us know what you think - if there are more recent / complete / credible studies out there which draw different conclusions, let us know!

    Engagement and participation

    The UK has a long history of public support for health research, as evidenced by the large number of participants in clinical trials and population studies (For example, the UK Collaborative Trial of Ovarian Cancer Screening and UK Biobank have recruited their targets of 200,000 and 500,000 individuals (respectively) with minimal objection to the use of their healthcare data) and the generous contributions to medical research charities such as Cancer Research UK and the British Heart Foundation.[1]

    Public engagement initiatives in relation to specific issues, such as the use of patient data, generally show that research is warmly supported. The attitudes of over 1,000 adults towards participating in health research were examined in the Wellcome Trust Monitor survey. Seventy-one per cent of participants indicated that they would be willing to give blood or tissue samples for research and 62% were willing to test a new treatment for a disease from which they were suffering.[2]

    Evidence from two national research studies demonstrates that a small number of patients complain about receiving direct invitations to participate in research. The UK Collaborative Trial of Ovarian Screening is one of the largest ever randomised controlled trials, covering 13 NHS Trusts in England, Wales and Northern Ireland, with successful recruitment of more than 200,000 women. Of the 1.2 million women invited to participate in the study only 32 complained about being contacted. UK Biobank reported from its integrated pilot phase that approximately 1 person from 1,000 invitations indicated that they did not want to participate because of concerns that their contact details had been provided to UK Biobank by the NHS.[3]

    There are a large number of organisations working to improve patient and public engagement with health research, including (but not limited to) UK Clinical Research Collaboration (UKCRC), INVOLVE, regulators themselves, the medical Royal Colleges, research charities and disease specific patient groups working to help the public understand the role and importance of research as an integral part of the care system. 

    Media view – data protection

    The influential role of mass media has important implications for the formation of public opinion and consequently public behaviour and the actions of policy makers. The most significant impact on attitudes towards the storage, transmission and use of personal data in healthcare is made by coverage of breaches of regulations and guidelines.

    Though many of the stories do not relate to data used in research per se, their impact contributes to patients’ concerns about any use of personal medical data.

    Storage and transmission of data are key to research, many large datasets (for example the national disease registries) inducting data from a variety of sources and releasing data for research to geographically dispersed users. A key aspect of the conduct of research is the ease with which those researchers can receive the data. The choice of transmission method is not driven solely by actual risk analysis: while an encrypted DVD has a high level of innate security, public perception of sensitive information being moved around on DVDs, memory sticks and laptops is an important consideration. In fact it was a major issue identified in the UK Ministry of Justice's report on Data Sharing, 2008.[4]

    A recent article from eWeekeurope.co.uk backs up perception with data under the inflammatory headline “A Freedom of Information request by… Software AG has revealed that most public sector bodies have no idea about secure data transfer.”[5] The article cites recent examples of the loss of sensitive information by public bodies: “A couple of years ago, Her Majesty’s Revenue and Customs (HMRC) lost a number of CDs containing private information on thousands of people. But there have been many more recent examples. Last July the UK Ministry of Defence admitted it had lost an entire server from a secure building – as well as 1.7 million individuals’ personal data. In November the UK Rural Payments Agency (RPA) lost backup tapes containing the payment and banking details of 100,000 farmers in the United Kingdom. And only last month an NHS worker in the secure mental health unit of a Scottish hospital was suspended, after he lost a USB stick containing patients’ medical records. The USB stick apprently contained unencrypted sensitive information – including the criminal histories of some violent patients at the Tryst Park unit at Bellsdyke psychiatric hospital. The stick was later found by a 12-year-old boy in the car park of an Asda supermarket.”

    The NHS was recently (April 2010) revealed by the Information Commissioner’s Office (ICO) to be responsible for the highest number of serious data breaches of any UK organisation since the end of 2007. David Smith, deputy commissioner at the ICO told the Infosec security conference the NHS had highlighted 287 breaches to it in the period, accounting for more than 30% of the total number reported.[6] Most of the breaches were the result of stolen data or hardware, followed by 82 cases of lost data or hardware. Richard Vautrey, the deputy chair of the British Medical Association's GPs committee thinks the number of breaches reflect the size and complexity of the NHS (the UK's largest employer with 1.7m staff) as well as its culture of openness.[7] Whilst comments in the BBC’s coverage mention in mitigation that the public sector’s culture of reporting all breaches contrasted with the private sector’s behaviour, these do little to lessen the impact of the headline: “NHS worst for data breaches.”


    [1] The Academy of Medical Sciences: A new pathway for the regulation and governance of health research
    [2] Ibid.
    [3] Ibid.
    [4] Ministry of Justice: Data Sharing Review, 2008 [Richard Thomas, Information Commissioner; Dr Mark Walport]
    [7] Ibid.

    Monday, February 6, 2012

    Data Sharing - the MRC Data Support Service and Research Data Gateway

    Whilst still working on an analysis of  research funders' policies with regard to data sharing we ought to update you on the progress made by the MRC on their Data Support Service. We posted on this in May 2011 and certainly  by late last year had not heard anything further but are now delighted to see that the MRC have a few new pages indicating that phase II ( "develop[ing] a prototype online gateway for the discovery of MRC-funded population and patient studies and their variables with a Directory of MRC population cohort datasets")  has completed and that phase III is underway about which you can read more on their site from whence these bullet points:

    • [Phase III is] Developing policy guidance for population and patient studies on sharing of research data and on data management planning, with expert input and in line with policies of other funder to ensure harmonised principles
    • Launching the prototype MRC Research Data Gateway to facilitate the discovery of research data, metadata and documentation
    • Planning and developing a sustainable MRC Research Data Gatewayand Directory of Population and Patient Research Data, with data management toolkit to support data sharing
    • Engaging further cohort studies to contribute metadata to the Directory of Population and Patient Research Data
    • Growing a data managers network with a programme of value-adding activities, to enable the preceding objectives














    We're particularly interested in the Research Data Gateway given our recent experience developing resource discovery portals (e.g. ONIX). The information available for each study accessible through the gateway is given below - we're not sure how to interpret 'Variable' and can't help but reflect on the lost opportunity of the now non-current DoCDat catalogue of clinical databases which provided very rich metadata on its resources including qualitative analyses - for more info on that see our blog post:

    • Study: a programme of research whereby data from and about individuals representing a population group are collected and analysed
    • Time period: a wave, sweep, time period or time point within a longitudinal study
    • Data collection event: a survey, screening, interview series or clinic, being the event through which research data were collected; there can be various data collection events within a phase
    • Variable: each data variable belongs to a particular study, phase and data collection event
    • Contact: point of contact for a study
    • Resource: an information item for a study, e.g. a questionnaire form, report, etc.


    Monday, January 30, 2012

    Data sharing in research - cultural and technical barriers in Life Sciences and Public Health - Part 2 - Barriers

    The second section of our review of attitudes to data sharing in Life Sciences and Public Health was due to  look at policies introduced by research funders to oblige researchers to make their data accessible to the community, however, there is more to be said on this than we have currently committed to bytes or paper.

    Thus we'll jump ahead here to our section on barriers to data-sharing and aim to get the section on funders' policies to you within the week...

    Just a reminder that it's meant to serve as a 'primer' aggregating analyses from the last few years in this area and pointing you to the original articles and as such references other publications pretty heavily - so do check out the footnotes if you want to explore further.

    Barriers – incentives and expertise

    Discussions of this subject highlight the lack of active incentives (rather than obligations) to share data: “the lack of explicit career rewards, and in particular the perceived failure of the Research Assessment Exercise (RAE) explicitly to recognise and reward the creating and sharing of datasets – as distinct from the publication of papers - are major disincentives.”[1]

    As mentioned above when looking at Social and Public Health Sciences the Research Intelligence Network found “found scant evidence of researchers wanting to publish datasets. Typically researchers will request data from one or more publicly-available datasets and they will undertake analysis. Often this process leads to the creation of new, derived datasets but these tend not to find their way to the public domain.”[2]

    Lack of incentive is blamed for this outcome: “Unlike in some of the other areas we have looked at there are no obvious rewards that accrue to researchers who decide to make their datasets publicly-available – though few deny that sharing datasets produced with public funds is a worthwhile principle….. researchers producing small scale datasets see no reason to invest the time and effort required to make their datasets publicly available. Besides which, some want to control their data, limit the possibility of the data being misrepresented, and limit the scope for competition [our italics].” [3]

    “Other disincentives include lack of time and resources; lack of experience and expertise in data management and in matters such as the provision of good metadata; legal and ethical constraints; lack of an appropriate archive service; and fear of exploitation or inappropriate use of the data….. Relatively few researchers have the expertise, resources and inclination to perform themselves all the tasks necessary to make their data not only available, but readily accessible and usable by others.”[4] Many researchers lack the skills to meet the quality standards imposed by data centres without substantial help from specialists.

    Additionally, “creating longitudinal datasets is an expensive business and therefore the people responsible for them tend to feel the need to protect them. This is manifested in reported anxiety about commercial organisations using data, deriving slightly or materially different datasets and claiming intellectual property rights over these new datasets.”[5]

    Across the biomedical sciences directors and PIs see their restricted data as their intellectual capital:  “As with most areas of research, there is competition between researchers to produce the best work in the best journal… Many researchers wish to retain exclusive use of the data they have created until they have extracted all the publication value they can.”[6]

    From the perspective of commercial / industry groups data sharing presents many of the same challenges: Intellectual Property and Confidentiality are particularly sensitive issues. In the past big pharmaceuticals organisations have traditionally been conservative over data sharing – concerns include loss of control, cost, other units reaching different conclusions or deriving novel insights which may have commercial value.

    Within this field there exists a significant heterogeneity of needs. The complexity and diversity of the biomedical research landscape breeds diversity in tools and methodologies for data capture and analysis, storage, maintenance and curation and this too may fuel confusion and dampen enthusiasm.


    [1] To Share or not to Share: Publication and Quality Assurance of Research Data Outputs - - Report commissioned by the Research Information Network (RIN) in association with the Joint Information Systems Committee and the National Environment Research Council (NERC) – published June 2008. This report covered six discrete research areas, two of which were Social and Public Health Sciences and Genomics and two interdisciplinary areas, one of which was Systems Biology.
    [2] Ibid.
    [3] Ibid.
    [4] Ibid.
    [5] Ibid.
    [6] Ibid.


    Monday, January 23, 2012

    Data sharing in research - cultural and technical barriers in Life Sciences and Public Health

    As promised for the PDF-shy - the first section of our review of attitudes to data sharing in Life Sciences and Public Health - it's also well worth reading the Research Information Network's (RIN) publication Data centres: their use, value and impact which RIN neatly summarise here. This review was meant to serve as a 'primer' aggregating analyses from the last few years in this area and as such references other publications pretty heavily - so do check out the footnotes if you want to explore further. We'll post the next section on Funders' policies tomorrow.

    Summary: Initiatives pursued by the funders and publishers of research to promote data-sharing have moved the agenda forward, however, cultural and technical barriers remain. Attitudes and abilities vary across specialities within biomedical research and though obliging data sharing has had some success, incentives and expertise are still lacking in many areas.



      The Current Situation

     The last ten years have seen consistent and targeted promotion of data sharing in research by organisations such as the Research Councils, the National Cancer Research Institute (NCRI), research charities and other funders and publishers of research. The aims are clear: “Ensuring data are made widely available to the research community accelerates the pace of discovery and enhances the efficiency of the research enterprise.”[1]

    The assumption is that data sharing is a ‘good thing’ and that connectivity between data will enable greater research potential. There has been a lot of activity in areas such as access and governance, initiatives providing portals and developing standards and there is evidence that “in many research fields – from genetics and molecular biology to the social sciences –data sharing is ingrained in how researchers work”[2] and that, with regard to Systems Biology at least, “data sharing, despite some anomalies, is the prevailing ethic.”[3]

    The picture is mixed across the life-sciences, however, as noted in this recent call for contributions to a thematic series on data standardisation, sharing and publication by the online Journal BM Research Notes: “different disciplines have embraced the possibilities of data sharing and open data to differing extents, and it can take the leadership of a small number of individuals to develop and promote their standard to secure widespread adoption, and enable interoperability of scientific data… In other cases a standard of data collection and preparation might be well known amongst circles of experts but perhaps unknown to researchers in different or even related fields. But with few journals considering data-driven articles and apparent inconsistencies in incentives and rewards for data publication, the availability of definitive and freely-available examples of re-usable, standardized data across the life sciences is patchy at best.”[4]

    Funders and journals are addressing this issue, promoting data sharing by various means including policies obliging researchers to make their output publicly available. However, “Where funder policies do not reach, there is a mix of results. Some researchers make great efforts to share data while others may retain their findings or publish in a form that means that although data are available, they are not readily accessible”[5] or ‘protecting by pdf’ as the practice is known within the community.

    The situation is least advanced in Social and Public Health Sciences: “There are many datasets produced by individual researchers or small project teams that could have long term viability if they were offered to an appropriate data centre, but this tends not to be the natural course of things. The sharing of datasets from small scale research projects appears to be relatively uncommon at present.”[6] Other analysts concur: “By contrast, this culture has yet to be widely embraced by the public health research community.”[7]

    Whilst the battle for cultural change appears to be far advanced in some areas and at least engaged in others, attitudes are not the only barrier: “Problems of reuse centre around… technical issues… – the variety of formats, the non-standardisation of formats, the need for proprietary software and so forth.”[8]

    The picture is not uniform across all fields within biomedical research - in some areas data sensitivity and cultural barriers remain more challenging to address than technical and ethical issues. Others, genomics for example, appear comparatively mature, with both cultural and technical issues well in hand.


    [1] Walport M, Brest P. Sharing research data to improve public health. The Lancet, Early Online Publication, 10 January 2011
    [2] Ibid.
    [3] To Share or not to Share: Publication and Quality Assurance of Research Data Outputs - Report commissioned by the Research Information Network (RIN) in association with the Joint Information Systems Committee and the National Environment Research Council (NERC) – published June 2008. This report covered six discrete research areas, two of which were Social and Public Health Sciences and Genomics and two interdisciplinary areas, one of which was Systems Biology.
    [4] A call for BMC Research Notes contributions promoting best practice in data standardization, sharing and publication; http://www.biomedcentral.com/1756-0500/3/235/
    [5] Ibid.
    [6] Ibid.
    [7] Walport M, Brest P. Sharing research data to improve public health. The Lancet, Early Online Publication, 10 January 2011
    [8] To Share or not to Share: Publication and Quality Assurance of Research Data Outputs