email us at Bioinformaticsolutions @ gmail . com

Showing posts with label research. Show all posts
Showing posts with label research. Show all posts

Monday, August 27, 2012

ELIXIR - 'A sustainable infrastructure for managing biological information in Europe' - One year on

It's been nearly a year since we last wrote about Elixir, the European research infrastructure project looking to support life-science information. Our friends at the European Bioinformatics Institute made us aware at the time that a pan-European project was under way to build and operate a sustainable infrastructure for managing and safeguarding biological information including genetic, protein and complex network analysis outputs.

So what's been happening in the intervening period?

Another seven countries have signed the Memorandum of Understanding in that time, broadening the remit and support-base for the initiative. The wider this base the better in light of the organisation's assertion that "the collection, curation, storage, archiving, integration and deployment of biomolecular data is an immense challenge that cannot be handled by a single organisation or by one country alone, but requires international coordination."

The European Commission's Community Research and Development Information Service (CORDIS) page on the project indicates that the first phase of funding is due to come to an end in December 2012 having run for five years. The project's aims over that period were all directed at gaining the widest possible support for the initiative by means of Memoranda of Understanding and included defining:
  • The scope of the infrastructure, its role and benefits
  • An appropriate governance and legal structure
  • A long term funding structure to provide a sustainable infrastructure
  • The requirements for the European Data Centre in the next 5-10 years
  • The critical interdisciplinary links that need to be forged between the biological and related scientific disciplines, including medicine, agriculture and the environment
  • The needs of related European industries
  • A training strategy to ensure that Europe effectively exploits all the available information
In carrying out this work Elixir assert that they were committed to involving all relevant stakeholders including users, data providers, tools providers to ensure that the infrastructure designed would be fit for purpose and exploring interoperability and supporting standards facilitating the between integration between core and specialised data resources.

Back in November 2011, we were anticipating the the Interim Board's announcement of the first phase of construction, however, we should point out that this period in Elixir's development is still the Preparatory Phase which title may make sense of the difficulty we've had finding concrete outputs from the project - the website for the Preparatory Phase is a little low-tech and also a little out of date - for more up-to-the-minute news see the Press Releases page on the main Elixir site from which you can see that much of the recent news involves the 'on-boarding' of various different European states but also covers the inception of their newsletter and, of most interest to us, the start of a new initiative co-ordinted by Elixir: BioMedBridges.

In their own words: "BioMedBridges is a joint effort of ten biomedical sciences research infrastructures on the ESFRI roadmap. Together, the project partners will develop the shared e-infrastructure—the technical bridges—to allow interoperability between data and services in the biological, medical, translational and clinical domains and thus strengthen biomedical resources in Europe. Launched in January 2012, the four-year initiative has been financed with €10.6 million by the European Commission’s Seventh Framework Programme."




Thursday, July 12, 2012

UK's e-infrastructure for research needs development

It may be too late to respond to the consultation (e-infrastructure for innovation and growth?) issued by the UK's research and education computer network JANET - but we'll be keeping an eye out for the results.

Their consultation was issued after JANET were asked to sit on the E-Infrastructure Leadership Council which has been created (in March 2012), co-chaired by David Willetts, Minister of State for Universities and Science, to co-ordinate the future governance and effective development of the UK’s research e-infrastructure.

It's worth checking out the consultation to see the kind of areas for development which JANET have in mind and for the link to the influential 2011 report 'A Strategic Vision for UK e-Infrastructure – A Roadmap for the Development and Use of Advanced Computing, Data and Networks'.

Thursday, June 21, 2012

Google BigQuery - Hospital Episode Statistics data analysis made easy (ish)

We had a visit from PA consulting last week, who were very excited about their application of Google's new Big Data solution, BigQuery, to the vast dataset which is the UK's Hospital Episode Statistics (HES).

We've had to wrestle with HES data before and come off the worse for it - and that was looking only at one year's Inpatient data - about 18 million lines of data. We encountered all kinds of issues in the setup of our bespoke database, created to hold and analyse this data, starting with the published data dictionary's divergence from the fields in our extract, taking in the discovery of duplicated unique episode identifiers (see our post on the response from the NHS Information Centre) and ending with our discovery that the processing power required to run some of our composite queries in a timely fashion was beyond our meagre infrastructure (this was at the now defunct National Cancer Research Institute's Informatics Initiative).

We could really have used the facility which BigQuery is set to provide. The guys from PA described how they had obtained the entire start-to-finish HES dataset across all three areas of collection (inpatient, outpatient and A&E) and loaded this into BigQuery (this being the most arduous part of the process, the data arriving on 27 DVDs and taking a couple of weeks to upload) prior to demonstrating the speed with which it was able to provide answers and how the data could be linked to google maps and google docs' spreadsheet application to dynamically produce visual and graphical analyses.

BigQuery dynamically calls in servers to assist in the running of a query based on the processing power required and then releases them once the query has been executed. The result of being able to access Google's immense army of servers is that without any of the usual time-consuming optimisation (indexing etc.) which supports enhanced performance on traditional database technologies, the user can execute a query against billions of data points in seconds. If you're working to some degree 'in the dark', uncertain of how you wish to structure your data and what analyses you will require it to support, you can experiment on vast datasets without waiting hours for queries to produce results (or fail!) - a facility which would have delivered huge time-savings to us in our HES analysis.

PA have a video describing their work and approach here and Google provide further information about BigQuery here.

Monday, May 14, 2012

Clinical Practice Research Datalink courts potential users

We received on Friday our invitation to the Clinical Practice Research Datalink (CPRD) Users Meeting that will be held at the MHRA offices in London on 24th May 2012, consisting of a series of short presentations by representatives of CPRD and external partner organisations to give "further insight, and an opportunity for input, into the current and future aims of CPRD."

The background we have covered severally before (see here and here): "CPRD is the new English NHS observational data and interventional research service, jointly funded by the NHS National Institute for Health Research (NIHR) and the Medicines and Healthcare products Regulatory Agency (MHRA). It combines the piloting work of the Research Capabilities Programme (RCP) and the existing General Practice Research Database (GPRD)."

CPRD services are designed to maximise the way anonymised NHS clinical data can be linked to enable many types of observational research and deliver research outputs that are beneficial to improving and safeguarding public health. CPRD will act to provide services to a wide range of researchers and the aim of the Users Meeting will be to ensure that its plans meet the needs of the broad cross section of researchers in academia, the NHS and commercial companies both in the UK and globally.

The specific topics that will be covered on the day include:
  • Pragmatic and Phase III - IV clinical trials
  • Multidimensional data quality
  • Hospital prescribing data
  • Models for linkage
  • Disease and patient group data marts
The day will provide "several opportunities for potential users to raise and discuss their priorities and requirements" to ensure CPRD meets researchers' needs.
It is possible that we won't be able to attend ourself so if any of our readers are planning on going along, let us know and we'll get in touch to see if you want to post some feedback on these pages!

Monday, May 7, 2012

Managing Research Data - the Joint Information Systems Committee Programme

In covering the BRISSkit vision for a cloud-based open-source "research application as a service" at the end of last year, we mentioned the Joint Information Systems Committee's (JISC)  Managing Research Data workstream about which we've been hearing a lot more recently and thought we should pass on the basics and a link for your own browsing:

JISC are a non departmental public body who support higher education and research in the UK by providing advice on the use of ICT - their Managing Research Data workstream supports both good data management and the sharing of data "for the benefit of UK Higher Education and Research". Their work in this area focuses on infrastructure, practice and skills:

  • piloting essential research data management infrastructures within institutions and for distributed research groups
  • improving practice in research data management planning
  • developing tools to help institutions plan their research data management practice
  • encouraging the publication of research data and demonstrating the benefits of improved methods for citing, linking and integrating research data
  • and, stimulating the acquisition of appropriate skills, among academics and research support staff in Universities
Follow this link to their page which contains further information about their internationally recognised Digital Curation Centre and more information on the five strands of the programme which include projects, planning, tools and training.

Monday, April 16, 2012

Oncology Research Information System now live at King's


We've blogged in the past about the Oncology Research Information System as an enterprise-breadth platform to support personalised cancer medicine.

When the System was presented at the NOCRI Information Systems Workshop towards the end of last year we gave you this overview taken from Prof. Peter Parker's abstract:

"ORIS is an IT platform that enables the routine extraction of consented, structured clinical data (notes, images, etc), its pseudonymisation and export into a research data store, from which patient cohorts can be selected and their information linked to molecular data derived from research"

We received an update from IDBS who are leading on the implementation and who look to have countered suggestions that the project was slow to deliver: "As you know until now, cross-hospital translational medicine technology to consolidate data across different scientific domains has been the subject of some debate but we have seen little movement in the industry in terms of actually making it happen. It is very pleasing to have the ORIS platform now in production [in] just under a year from when we started... with a lot of work that is now in the product so as we implement now the time scales are much reduced."

They also sent us a link to the press-release which gives further information and includes a link to details of their Enterprise Translational Medicine Solution.

Wednesday, April 4, 2012

Clinical Practice Research Datalink is finally here - or is it?

The new Clinical Practice Research Datalink about which we have blogged much in the past has finally arrived (http://www.cprd.com/intro.asp) amid a certain amount of fanfare - see this from Pharma Times, this from PMLive and this from GP magazine.

 You may notice that the GPRD pages now redirect to this site and to a certain extent, this is largely a rebranding exercise at the moment. Behind the scenes a team at the DoH are trying to ensure that major data sources are willing and able to engage with this initiative but the speed at which they come online remains to be seen. We'll fill you in further on plans for a researchers data-catalogue interface as this project advances - and if CPRD are not offering that just yet, perhaps the MRC are - we'll get you up to date with the MRC's Data Support Service before the week is out.



Monday, February 27, 2012

News from the EMBL-European Bioinformatics Institute: Enzyme Portal Launches


We received the following press release today from the European Bioinformatics Institute announcing the launch of their Enzyme Portal, a freely available resource for people who are interested in the biology of enzymes and proteins with enzymatic activity:

"Enzymes catalyse the myriad reactions that take place in living organisms, allowing chemical changes to occur that would otherwise need conditions that are incompatible with life. Until now, information about enzymes was scattered throughout many different resources. This meant that you had to know exactly what you were looking for, which could be a real barrier to discovery.

The Enzyme Portal mines and displays data about proteins with enzymatic activity from public repositories via a single search, and includes biochemical reactions, biological pathways, small molecule chemistry, disease information, 3D protein structures and relevant scientific literature. It summarises information in the UniProt knowledge base; the Protein Data Bank in Europe; Rhea, a database of enzyme-catalysed reactions; Reactome, a database of biochemical pathways; IntEnz, a resource with enzyme nomenclature information; ChEBI and ChEMBL, which contain information about small-molecule chemistry and bioactivity; and CoFactor and MACIE for highly detailed, curated information about cofactors and reaction mechanisms.

The Enzyme Portal covers a large number of species, including the key organisms used in biological research, and makes it simple to compare the characteristics of equivalent enzyme activities in different organisms.

“The Enzyme Portal will serve researchers interested in the native metabolism of organisms as well as those working in drug discovery or chemical biology,” explained Dr Christoph Steinbeck, Head of Cheminformatics and Metabolism at EMBL-EBI. “The resource seamlessly bridges the various enzyme-related aspects of these areas, all the way from the small molecules – which may occur naturally in the organism of study or be introduced  – to the 3D structure of the enzymes affected and the genomic information coding for those.”

The design of the Enzyme Portal was based entirely on user demand and feedback. “We didn’t come in with preconceived ideas of what it should be; it was designed by users, for users,” said Cheminformatics Coordinator Paula de Matos. “We had enough data, but it was distributed over 10 different resources. Now, we have a central place where they can access and explore all of it – including information about disease.”

“The Enzyme Portal takes the researcher’s perspective,” added User Experience Analyst Jenny Cham. “We leveraged user-testing feedback to create it, making it the first EMBL-EBI resource to have a fully user-centred design from scratch. This approach was not only cost-effective; it made decision making and communication much easier – particularly in terms of design and technology choices.”



To watch a short video about the Enzyme Portal on our YouTube channel, visit: http://youtu.be/Kldp0WXcxUM

Wednesday, February 22, 2012

General Practice data for research

We have just read, courtesy of eHealth Insider that "the NHS Information Centre is on the verge of having all GP clinical systems suppliers signed up to the General Practice Extraction Service. Emis, Microtest, iSoft and INPS have signed up to extract and communicate data to the NHS IC, and TPP expects to have a contract signed within weeks."


For more information on the General Practice Extraction Service (GPES), see the NHS IC GPES page which currently only mentions the contract with EMIS and whilst it describes the value of GP data, does not allude to its use for research but only in commissioning despite asserting that "GP patient records are the most complete record of a patient's health within the NHS. They comprise a wealth of information about patient care, the prevalence of diseases and treatments given."

So where does this leave services like GPRD, THIN and QRESEARCH. Well, we know that GPES won't be fully up and running for another year at least, and that GPRD will become the CPRD  (Clinical Practice Research Datalink) if all goes to plan. In the interim, as per this from the Health Protection Agency, GPRD and QRESEARCH will still be providing primary care data for research, but what will the landscape look like by the end of 2013?

Tuesday, February 21, 2012

Patients' views on participating in medical research - Part 2: Engagement and Participation

Our last post here began our summary of recent work on patients' attitudes to participating in medical research - part of our primer on the current state of play in clinical research in the UK selecting choice elements from recent reports; check out the footnotes for interesting sources to follow up. We include in this a look at the influence of media coverage of issues pertaining to data security on patients' attitudes and behaviour.

As always, let us know what you think - if there are more recent / complete / credible studies out there which draw different conclusions, let us know!

Engagement and participation

The UK has a long history of public support for health research, as evidenced by the large number of participants in clinical trials and population studies (For example, the UK Collaborative Trial of Ovarian Cancer Screening and UK Biobank have recruited their targets of 200,000 and 500,000 individuals (respectively) with minimal objection to the use of their healthcare data) and the generous contributions to medical research charities such as Cancer Research UK and the British Heart Foundation.[1]

Public engagement initiatives in relation to specific issues, such as the use of patient data, generally show that research is warmly supported. The attitudes of over 1,000 adults towards participating in health research were examined in the Wellcome Trust Monitor survey. Seventy-one per cent of participants indicated that they would be willing to give blood or tissue samples for research and 62% were willing to test a new treatment for a disease from which they were suffering.[2]

Evidence from two national research studies demonstrates that a small number of patients complain about receiving direct invitations to participate in research. The UK Collaborative Trial of Ovarian Screening is one of the largest ever randomised controlled trials, covering 13 NHS Trusts in England, Wales and Northern Ireland, with successful recruitment of more than 200,000 women. Of the 1.2 million women invited to participate in the study only 32 complained about being contacted. UK Biobank reported from its integrated pilot phase that approximately 1 person from 1,000 invitations indicated that they did not want to participate because of concerns that their contact details had been provided to UK Biobank by the NHS.[3]

There are a large number of organisations working to improve patient and public engagement with health research, including (but not limited to) UK Clinical Research Collaboration (UKCRC), INVOLVE, regulators themselves, the medical Royal Colleges, research charities and disease specific patient groups working to help the public understand the role and importance of research as an integral part of the care system. 

Media view – data protection

The influential role of mass media has important implications for the formation of public opinion and consequently public behaviour and the actions of policy makers. The most significant impact on attitudes towards the storage, transmission and use of personal data in healthcare is made by coverage of breaches of regulations and guidelines.

Though many of the stories do not relate to data used in research per se, their impact contributes to patients’ concerns about any use of personal medical data.

Storage and transmission of data are key to research, many large datasets (for example the national disease registries) inducting data from a variety of sources and releasing data for research to geographically dispersed users. A key aspect of the conduct of research is the ease with which those researchers can receive the data. The choice of transmission method is not driven solely by actual risk analysis: while an encrypted DVD has a high level of innate security, public perception of sensitive information being moved around on DVDs, memory sticks and laptops is an important consideration. In fact it was a major issue identified in the UK Ministry of Justice's report on Data Sharing, 2008.[4]

A recent article from eWeekeurope.co.uk backs up perception with data under the inflammatory headline “A Freedom of Information request by… Software AG has revealed that most public sector bodies have no idea about secure data transfer.”[5] The article cites recent examples of the loss of sensitive information by public bodies: “A couple of years ago, Her Majesty’s Revenue and Customs (HMRC) lost a number of CDs containing private information on thousands of people. But there have been many more recent examples. Last July the UK Ministry of Defence admitted it had lost an entire server from a secure building – as well as 1.7 million individuals’ personal data. In November the UK Rural Payments Agency (RPA) lost backup tapes containing the payment and banking details of 100,000 farmers in the United Kingdom. And only last month an NHS worker in the secure mental health unit of a Scottish hospital was suspended, after he lost a USB stick containing patients’ medical records. The USB stick apprently contained unencrypted sensitive information – including the criminal histories of some violent patients at the Tryst Park unit at Bellsdyke psychiatric hospital. The stick was later found by a 12-year-old boy in the car park of an Asda supermarket.”

The NHS was recently (April 2010) revealed by the Information Commissioner’s Office (ICO) to be responsible for the highest number of serious data breaches of any UK organisation since the end of 2007. David Smith, deputy commissioner at the ICO told the Infosec security conference the NHS had highlighted 287 breaches to it in the period, accounting for more than 30% of the total number reported.[6] Most of the breaches were the result of stolen data or hardware, followed by 82 cases of lost data or hardware. Richard Vautrey, the deputy chair of the British Medical Association's GPs committee thinks the number of breaches reflect the size and complexity of the NHS (the UK's largest employer with 1.7m staff) as well as its culture of openness.[7] Whilst comments in the BBC’s coverage mention in mitigation that the public sector’s culture of reporting all breaches contrasted with the private sector’s behaviour, these do little to lessen the impact of the headline: “NHS worst for data breaches.”


[1] The Academy of Medical Sciences: A new pathway for the regulation and governance of health research
[2] Ibid.
[3] Ibid.
[4] Ministry of Justice: Data Sharing Review, 2008 [Richard Thomas, Information Commissioner; Dr Mark Walport]
[7] Ibid.

Monday, February 6, 2012

Data Sharing - the MRC Data Support Service and Research Data Gateway

Whilst still working on an analysis of  research funders' policies with regard to data sharing we ought to update you on the progress made by the MRC on their Data Support Service. We posted on this in May 2011 and certainly  by late last year had not heard anything further but are now delighted to see that the MRC have a few new pages indicating that phase II ( "develop[ing] a prototype online gateway for the discovery of MRC-funded population and patient studies and their variables with a Directory of MRC population cohort datasets")  has completed and that phase III is underway about which you can read more on their site from whence these bullet points:

  • [Phase III is] Developing policy guidance for population and patient studies on sharing of research data and on data management planning, with expert input and in line with policies of other funder to ensure harmonised principles
  • Launching the prototype MRC Research Data Gateway to facilitate the discovery of research data, metadata and documentation
  • Planning and developing a sustainable MRC Research Data Gatewayand Directory of Population and Patient Research Data, with data management toolkit to support data sharing
  • Engaging further cohort studies to contribute metadata to the Directory of Population and Patient Research Data
  • Growing a data managers network with a programme of value-adding activities, to enable the preceding objectives














We're particularly interested in the Research Data Gateway given our recent experience developing resource discovery portals (e.g. ONIX). The information available for each study accessible through the gateway is given below - we're not sure how to interpret 'Variable' and can't help but reflect on the lost opportunity of the now non-current DoCDat catalogue of clinical databases which provided very rich metadata on its resources including qualitative analyses - for more info on that see our blog post:

  • Study: a programme of research whereby data from and about individuals representing a population group are collected and analysed
  • Time period: a wave, sweep, time period or time point within a longitudinal study
  • Data collection event: a survey, screening, interview series or clinic, being the event through which research data were collected; there can be various data collection events within a phase
  • Variable: each data variable belongs to a particular study, phase and data collection event
  • Contact: point of contact for a study
  • Resource: an information item for a study, e.g. a questionnaire form, report, etc.


Monday, January 30, 2012

Data sharing in research - cultural and technical barriers in Life Sciences and Public Health - Part 2 - Barriers

The second section of our review of attitudes to data sharing in Life Sciences and Public Health was due to  look at policies introduced by research funders to oblige researchers to make their data accessible to the community, however, there is more to be said on this than we have currently committed to bytes or paper.

Thus we'll jump ahead here to our section on barriers to data-sharing and aim to get the section on funders' policies to you within the week...

Just a reminder that it's meant to serve as a 'primer' aggregating analyses from the last few years in this area and pointing you to the original articles and as such references other publications pretty heavily - so do check out the footnotes if you want to explore further.

Barriers – incentives and expertise

Discussions of this subject highlight the lack of active incentives (rather than obligations) to share data: “the lack of explicit career rewards, and in particular the perceived failure of the Research Assessment Exercise (RAE) explicitly to recognise and reward the creating and sharing of datasets – as distinct from the publication of papers - are major disincentives.”[1]

As mentioned above when looking at Social and Public Health Sciences the Research Intelligence Network found “found scant evidence of researchers wanting to publish datasets. Typically researchers will request data from one or more publicly-available datasets and they will undertake analysis. Often this process leads to the creation of new, derived datasets but these tend not to find their way to the public domain.”[2]

Lack of incentive is blamed for this outcome: “Unlike in some of the other areas we have looked at there are no obvious rewards that accrue to researchers who decide to make their datasets publicly-available – though few deny that sharing datasets produced with public funds is a worthwhile principle….. researchers producing small scale datasets see no reason to invest the time and effort required to make their datasets publicly available. Besides which, some want to control their data, limit the possibility of the data being misrepresented, and limit the scope for competition [our italics].” [3]

“Other disincentives include lack of time and resources; lack of experience and expertise in data management and in matters such as the provision of good metadata; legal and ethical constraints; lack of an appropriate archive service; and fear of exploitation or inappropriate use of the data….. Relatively few researchers have the expertise, resources and inclination to perform themselves all the tasks necessary to make their data not only available, but readily accessible and usable by others.”[4] Many researchers lack the skills to meet the quality standards imposed by data centres without substantial help from specialists.

Additionally, “creating longitudinal datasets is an expensive business and therefore the people responsible for them tend to feel the need to protect them. This is manifested in reported anxiety about commercial organisations using data, deriving slightly or materially different datasets and claiming intellectual property rights over these new datasets.”[5]

Across the biomedical sciences directors and PIs see their restricted data as their intellectual capital:  “As with most areas of research, there is competition between researchers to produce the best work in the best journal… Many researchers wish to retain exclusive use of the data they have created until they have extracted all the publication value they can.”[6]

From the perspective of commercial / industry groups data sharing presents many of the same challenges: Intellectual Property and Confidentiality are particularly sensitive issues. In the past big pharmaceuticals organisations have traditionally been conservative over data sharing – concerns include loss of control, cost, other units reaching different conclusions or deriving novel insights which may have commercial value.

Within this field there exists a significant heterogeneity of needs. The complexity and diversity of the biomedical research landscape breeds diversity in tools and methodologies for data capture and analysis, storage, maintenance and curation and this too may fuel confusion and dampen enthusiasm.


[1] To Share or not to Share: Publication and Quality Assurance of Research Data Outputs - - Report commissioned by the Research Information Network (RIN) in association with the Joint Information Systems Committee and the National Environment Research Council (NERC) – published June 2008. This report covered six discrete research areas, two of which were Social and Public Health Sciences and Genomics and two interdisciplinary areas, one of which was Systems Biology.
[2] Ibid.
[3] Ibid.
[4] Ibid.
[5] Ibid.
[6] Ibid.


Wednesday, January 18, 2012

Attitudes to Data Sharing in Research

Thanks to the NHS-Higher Education Forum for publishing some of the work we've done for the National Cancer Research Institute - namely a review of attitudes to data sharing in research - as part of the output of their NHS-HE Connectivity Working Group's Information Governance and Data Sharing workstream.

We'll reproduce the content as a series on this blog over the next few days for the PDF-shy.

Monday, January 9, 2012

Response to the National Cancer Intelligence Network's consultation: Building an e-health research infrastructure for cancer

Today is deadline day for responses to the National Cancer Intelligence Network's (NCIN) consultation: Building an e-health research infrastructure for cancer (see our blog entry for a list of their proposals).

We have submitted a response from the perspective of those who have worked with data supporting basic and translational research - we'll also be posting later a link to the NCIN and Department of health publication from December last year: An Intelligence Framework for Cancer - but for now, here's how we replied to the consultation:



The UK is currently supporting or considering the development of several initiatives seeking to promote the skills and infrastructure necessary to carry out health research based on linked large-scale or population-level datasets generated through routine processes of data collection.

In Wales, the Health Information Research Unit at the University of Swansea maintains the Secure Anonymised Linkage System (SAIL); in Scotland, the Scottish Health Informatics Programme (SHIP) supports the “collation, management, dissemination and research analysis of anonymised Electronic Patient Records”; in the UK, the Research Capability Programme of the National Institute for Health Research have piloted a Health Research Support Service which is due to be formally implemented as a full service:  the Clinical Practice Research Datalink. The Medical Research Council also recently issued a call for e-Health Informatics Research Centres to “maximise the health research potential offered by linking electronic health records with other forms of routinely collected data and research datasets”.

Of the National Cancer Research Institute partners’ 2010 funding, however, over 50% was spent on research which could be described as basic, translational or early stage; 40% on Biology and Aetiology alone; and some proportion of the discovery and development elements of spending under Common Scientific Outcome (CSO) 5 (Treatment), technology development and evaluation under CSO 4 (Early Detection, Diagnosis and Prognosis), and CSO7 (Scientific and Model Systems) can be ascribed to these types of research.

The infrastructure used to support this work is in many cases intra-institutional, in some inter-institutional, rather than national – although with appropriate standardization, integrated datasets from within institutions could be submitted to national-scale repositories with greater ease. Completeness, accuracy and granularity of the data are vital for this research. Often the data which support and contextualize observations in the laboratory during these research projects are drawn from multiple hospital systems and collated with difficulty. This has an impact on timescales and the validation of observations. Some proportion of the NCRI partners’ spend in each of the CSOs is dedicated to Resources and Infrastructure (R&I) which may include informatics, however, if we look at the other project types falling under R&I for e.g. CSO 4 (Early Detection, Diagnosis and Prognosis) we may reasonably conclude that the proportion dedicated to informatics is not the majority – closer analysis of the NCRI CaRD database is required to confirm this.
CSO4.4 Examples of science that would fit:

·         Informatics and informatics networks; for example, patient databanks
·         Specimen resources (serum, tissue, images, etc.)
·         Clinical trials infrastructure
·         Epidemiological resources pertaining to risk assessment, detection, diagnosis, or prognosis
·         Statistical methodology or biostatistical methods
·         Centers, consortia, and/or networks
·         Education and training of investigators at all levels (including clinicians), such as participation in training workshops, advanced research technique courses, and Master's course attendance. This does not include longer term research based training, such as Ph.D. or post-doctoral fellowships

The MRC, in their call for e-Health Informatics Research Centres, adduce the key findings of the ABPI and UK research funders mapping exercise reviewing the UK capability in e-Health records research – a number of these can be applied to the intra-institutional situation:: institutions could be submitted to national-scale repositories with greater easeics is not the
·         There is a shortage of people with the breadth of skills necessary to carry out the complex linkage and analyses required in health informatics research.
·         There is an absence of career structure in enabling roles such as data managers, software engineers, informaticians and data analysts.
·         There are no clear interfaces between researchers and industry, policy makers or the NHS and there is no ready means for sharing best practice.

Certainly, my own experience of supporting even institutes with strong reputations for research is that they lack the skills, focus and confidence in informatics to make much progress in the development of their infrastructure – and have been extremely glad of the opportunity to take advice and receive support from experienced individuals with a research and informatics background.

Perhaps the NCIN could consider devoting some resource to skills development in this area, disseminating the acquired expertise and knowledge of the NCIN of best practices in data management and handling and the use of technology. Might this sit alongside the work currently envisaged by Proposal 6 of the consultation?
_________

During a meeting with Oracle at the end of last year, an ex-colleague who specialises in molecular and gynaecological oncology suggested that their institute would not be seeking data integration services and infrastructure supply from the likes of Oracle with such urgency if they felt they could get ‘stage and grade’ at diagnosis from the Thames Cancer Registry.

The paucity of staging data in the registries is an established weakness as discussed in the NCIN and Department of Health document, An Intelligence Framework for cancer and steps are being taken to address this, however, the perception of the inadequacy of the dataset collected by the Thames Cancer Registry (and by extension, despite shining examples such as the ECRIC, the amalgamated registries’ dataset) extended beyond the known weaknesses unfairly to the dataset as a whole in the case of this Professor. Such perceptions were not uncommon at that centre and need to be overturned.

The vastly extended dataset which will be collected by the registries in future sounds extremely promising in its potential to support not only epidemiological and population-level research, but also basic and small-scale clinical research. It will be vital, however, to create a sustained ‘sales’ initiative to establish a new level of confidence in the data in the areas of the research community who have hitherto not engaged with these datasets due to the concerns described in An Intelligence Framework. Their concern may be that where a smaller dataset was found wanting, will the collection of a larger one not push already stretched resources beyond their elastic limit?

Having had first-hand experience of the way in which MDT data is fed into the Somerset system and the ample opportunities, often taken by overburdened MDT co-ordinators, to introduce error – it is inspiring to see that a truly modern approach to data extraction and aggregation is being implemented as described by Dr. Rashbass at, to give one instance, the NOCRI Information Systems Workshop.  As described by Dr. Rashbass, various technologies including natural language querying will take data from pathology full-text reports, from local imaging systems and myriad other systems to create the amalgamated national dataset – and this data will be quality controlled and assured. More information on how the latter will be achieved would be welcome.

Similar initiatives and technologies are being employed by healthcare delivery and research organisations themselves – for example, the ORIS oncology platform being implemented intra-organisationally by King’s Health Partners and the Acropolis platform being implemented inter-organisationally. It is important to note that these implementations may be beyond the budget of smaller organisations who deliver oncology services and conduct research – and here the value of a new ‘high-resolution’, quality assured, timely dataset such as that envisaged by the registry modernisation team will have the potential to deliver enormous benefit.

But this will depend on the quality of the data and ensuring that this quality is recognised in the research community. “This service will ensure that common standards and working practices are applied to data extraction, linkage and quality assurance to both national feeds and a range of local sources.” This assertion really needs to be backed up with a strong communications and ‘marketing’ effort.

To this end, should the NCIN devote some resource to support activities at the provider end of the process to ensure that where providers are implementing their own data infrastructures, these can interface with and provide bulk data to the unified registries to the appropriate standard; and where they are not yet capable of developing their own infrastructures, that they have support in the provision of accurate and complete data to the registries and potentially support in the process of designing their own data architectures and integration solutions; and then to effectively communicate the work that they are doing to improve registry data effectively to the community – concentrating not on the sophisticated use of technology to capture and amalgamate data, but on the procedural changes being implemented to assure quality?

Many of the proposals made in the consultation document might be realised by the same infrastructural components – and many of these components are similar to those which will hopefully be implemented by the Clinical Practice Research Datalink (CPRD). Where respondents to the consultation indicate that the proposed data linkage and notification services would deliver great benefit to their work, it may be worth establishing what level of awareness they have of the CPRD, the concern being that the overhead involved in creating facilities which might duplicate some aspects of the CPRD could be enormous given the proportion of the budget for the latter initiative devoted to infrastructure. There might be a greater return on investment to be had by focusing on data rather than infrastructure at the national level?

In conclusion, it might be worth considering if Proposal 6 (a research support service advising on the availability of and access to data) could benefit from being expanded to include some work looking at supporting data quality and intra-institutional infrastructural development -  and engaging the basic and translational research communities to overturn perceptions about the ability of the dataset to support their work.


Let us know what you think - are we way off-beam?

Wednesday, January 4, 2012

Data Integration platforms - Research and Service Delivery - Oracle

We're back to full strength now and looking forward to passing more in-depth info your way this year than hitherto - starting with a look at data integration platforms (Oracle, Orion, IBM, Cerner, OpenClinica and more) being considered for use in UK and US healthcare institutions looking to harness the power of their currently fragmented data to deliver improved services and support research.

We recently started working with a major London-based cancer research centre who are entertaining pitches from various suppliers looking to pull together their pathology, cytogenetics, radiology and clinical data (plus everything else, immunology, toxicology, virology, you name it) and we've been digging into the detail behind the glossy slides - starting with Oracle:



Oracle have an increasing presence in this space in the US and are looking to use their learning there to expand their healthcare division in the UK. Their current UK consulting workforce (healthcare) is still in the single digits, but they sought to impress with details of their R&D spend and some recent examples of their work in the States. For some interesting discussion around their presentation of R&D spend, see these two bloggers (enterpriseirregulars and martijnlinssen) who discuss absolute spend vs. proportional spend and what it means about Oracle's R&D budget compared to, for example SAP.

Key to their presentation was a slide describing the systems architecture of a completed end-to-end solution with the Oracle Health Data Warehouse Foundation at the core and an 'Omics analysis platform linked in which we thought was pitched as the value-adding component.

Certainly the 'Omics platform is indicative of their aim to get to work in the translation medicine area (described as the North Star of their current thrust) - but it's an undeveloped product, not even fully implemented at Moffitt as far as we are aware. And that will be its first implementation - Moffitt hinted that this was the missing component in the Oracle solution the first time round (the guys from Oracle, however, countered that at least their core integration engine worked where Microsoft and Orion had both previously faltered).

See here for an interesting article which describes how a number of other vendors may be brought in to perform analytics on the core Oracle database.

As it happens, Oracle are also developing a Translation Research Center product - see here for Oracle's full healthcare product offering.

We're familiar with some of the Oracle products already - the Health Transaction Base, it's Enterprise Terminology Services - the latter is in use at at King's Health in a their enterprise oncology solution being implemented by IDBS (see our blog entry) - but King's Health are using ORION's Rhapsody data integration product rather than an Oracle solution - we will try to get answers from Peter Parker as to their rationale.

We had a few issues with their presentation which centred on that solution architecture slide which featured the Oracle solutions centre-stage in bright red, and off to the left in grey was the data extraction process from the multiple systems currently in place in any healthcare provision organisation. They admitted that this small grey box is where the majority of the work and cost is incurred!

Of note, however, they are developing a series of app-style products which can then plug into the clinical and genomic data warehouses to fulfil various research and service delivery ends - looks interesting and we will try to get you more info on this.

As imagined, it looks like the major pain point in any implementation is extraction from clinical systems such as those Cerner and others are currently managing in the UK - we're going to be in touch with some folks who have developed bespoke SQL products to perform exactly that task in the coming weeks - keep checking here for more on this - and we'll take a look at some of the other products in this 'space' - so watch this [space].

For a discussion of the Dana Farber's experience implementing Oracle - see here.

Wednesday, November 30, 2011

"Newly available Health Data will support Medical Research and Patient Empowerment"

That's the line from the Department of Health! We received a link to this in response to a query we submitted questioning the accuracy of a report in the Guardian which contained the following paragraph:

"The health data will be the most comprehensive available outside of US veterans' medical records in America, publishing anonymous records of medical treatment from GP to hospital –something proposed by the Wellcome Trust in its submission to the NHS Innovation review this year: "Integrated databases … would make England unique, globally, for such research." Medical researchers and big pharmaceutical companies will be able to use the data for free"


It's a piece of poor reporting which slightly beggars belief in the light of the Guardian's coverage yesterday of the Leveson inquiry and revelations (or confirmations, rather) of shabby reporting practices by the tabloids. They have either poorly researched, poorly understood or deliberately conflated items in order to inspire the wrath of Guardian commentators who responded in part by lamenting, as they see it, the government's gift to pharmaeutical companies eager to use this data to develop new drugs to sell to the NHS at inflated prices - I'm thinking of this comment in particular:


"private pharamaceutical companies to analyse health data in order to make more money out of health for the private sector. Meanwhile every week a new study finds something "wrong" with the NHS -manufacturing consent that what the country really needs is private healthcare.
Really loving the fact that my health records will be analysed so that private companies can prosper and develop "remedies" that will be sold to hospitals for a fortune.
We are witnessing a sick and hopefully dying civilization."
From what we understand, there is a conflation here of the NHS Information Centre's decision to release prescription data and the previously blogged about Clinical Practice Research Datalink which will not be releasing information for free - check out the release from the DoH for clarity and for some excellent commentary see Simon Denegri's blog and Becky's Policy Pages.