Pages

Showing posts with label impact of grids. Show all posts
Showing posts with label impact of grids. Show all posts

Wednesday, September 19, 2012

Proving, Recognising and Measuring Impact

Over the last two days of the EGI-Technical Forum, we've been extremely busy interviewing speakers, but I have managed to dip into a couple of sessions related to the challenge of measuring impact.

Yesterday's session on Scientific Publications Repository workshop highlighted the increasingly weighty issue of where and how to store scientific outputs (i.e. papers and associated data) from EU projects.  The important subject of acknowledgement of e-infrastructures was realised. It is apparent that while many researchers depend on increasing computational capacity most would not be able to identify the provider.  Scientists are interested in disseminating their achievements so while they might refer to a supercomputer or grid when publishing their results they generally do not specify the exact nature of the resources. But the fact remains if the end users of grids are unaware, then it is impossible for others to recognise the contribution these resources are making to scientific literature. With EC policies now in place to develop Open Access to research results, the session provided a perfect opportunity to explore the best ways to provide access and to appropriately acknowledge these resources.

Fortunately, there are projects examining the issue of harvesting this information. During the session, we heard from Pasquale Pagano, who described the OpenAIRE plus project. Their project is collecting and linking EU publications in an open access repository, which will give researchers both access to research material and valuable statistics to measure their impact.  Pagano gave us a glimpse into data that will be available e.g. publications per project, statistics for areas/institutions, publications (see slides). There are already 26,000 FP-7 publications from 5,000 projects (>10K Open Access). How this fits in at a national, international level was also discussed.

In these recessionary times, there is understandably an increasing pressure to prove your project's efficiency and effectiveness. The ERINA+ project have been working on developing tools for evaluating the scientific impact of EU-funded infrastructures. They shared their preliminary results with us this afternoon at the conference. Take it from me, tasked with measuring our own project's impact, it is a real challenge to capture all the predicted and unprecedented outcomes of our own dissemination project. It has been instructive learning about their methodology, and seeing the  ERINA+ project standardise their self-evaluation web tool for 21 projects.

So far, the aggregated (preliminary) results show that these projects have collectively conducted 213 dissemination activities reaching out to 246,000 persons in Europe and outside including United States, Latin America and China. Four spin-offs have resulted (no details provided in the slides). The projects have also been quite prolific! In total, 114 papers with impact factor have been published, and a further 62 papers without impact factor (plus 96 articles in conference proceedings). ERINA+ will provide a personalised report for projects later in the year, and will presenting their findings at eChallenges in October.

Watch our interview with Andrea Manieri (ERINA+ project coordinator) to find out more.


Tuesday, September 15, 2009

Biomedical Informatics without Borders: Chatting with Martin Hofmann-Apitius

An aside from all the EGEE'09 prep going on in our office, here's my interview with Martin Hofmann-Apitius from the Fraunhofer Institute, who chaired HealthGrid 2009 earlier this year. Martin will soon be appearing in our GridBriefing on eHealth – I caught up with him at the 'Biomedical Informatics without Borders' conference last week.

Do you think eHealth and grid computing has a lot to offer healthcare?

I think [grid computing for health] has a lot of potential. What we have to keep in mind, is that we have to be prepared to go a long, long way with it. It was the grid three years ago – now everybody's talking about the cloud, and actually I don't really care because the point is about interoperability. It's about working together, it's about all the e-science paradigms.

eHealth takes much longer then in the high energy physics community (HEP) because in HEP you have a couple of hundred people working there. In health you have hundreds of thousands and in eHealth you still have a lot of people, much more than in HEP. Therefore finding a common understanding is more difficult in that community and there's all these privacy, and legal and ethical considerations which make things much more complicated. When you look at HEP, the particles that they're looking at they have no will. So they cannot declare their privacy rights - but humans are really complex. And they claim 'I don't want to share my data, I don't want to submit my genome'. Craig Venter has got his genome sequence and made it publicly available but other people are scared of doing that, and that means we've got to be patient with grids and with eHealth.

They have big potential but it will take much longer than the engineering disciplines, the high energy physics, the meterology community, the astronomy community and so on.

If it's going to be a long journey do you think it's important to invest more time, or money or for grid researchers and technicians to co-operate more closely with physicians and clinicians?

There's the human factor – mutual respect and understanding even for opinions that are in contradiction to the eScience paradigm, this is one thing that's necessary. I don't necessarily believe that more money would help; I think that the funding regimes have to be sustainable. In the eHealth arena we have to be aware that impact assessment – the impact of caBIG and other grid health biomedical research infrastructures on health comes in ten years or twenty years. Not in 2010, not in 2012, maybe in 2020. This is something where I'm actually in alignment with Otto Rienhoff (the head of the German MediGRID project). He always says we must not oversell, we must be careful communicating that the grid is the clue to all the problems people have in the healthcare systems.

We will not, in a short time, improve cancer treatment, we will not reduce costs of the healthcare system. We do research on how to use distributed computing, distributed data management and shared semantics to create problem solving environments which ultimately, in a couple of years, will have an effect on costs in the healthcare system or the frequency of discovering new insight - but currently we're still doing groundwork.

So it's a bit of a cautionary tale then?

We should just refrain from producing hype. We should be careful in communication of promises for the people who ultimately pay our salary - the taxpayers - and people in politics, we really have to be careful with what we promise.

On the other side when I look at what people such as Carole Goble [the director of myGrid] and caBIG are doing, it's quite impressive to see what advancements have been made there and how things are evolving so on the other side I'm quite optimistic from a technology point of view.

Friday, September 11, 2009

Biomedical Informatics without Borders: Talking to Ken Buetow

While collecting some quotes to use in our upcoming GridBriefing on eHealth I was lucky enough to chat to Ken Buetow, associate director at the National Cancer Institute in the US and leader of caBIG (cancer Biomedical Informatics Grid). Here's what he had to say.

What do you think grids and bioinformatics can offer healthcare?


I think that they're transformative technologies. I think they represent an opportunity to connect the disparate, and I might say at times desperate, components of health and healthcare, medicine, biomedicine and biomedical research in a manner that the whole can become more than the sum of the parts. [They] allow us to leverage the observations [...] in a healthcare or healthcare delivery setting, to help us know what the important research questions are – what's working, what's not working - and to bring the information into a research setting where we can discover new ways to approach disease, new interventions to treat disease and to understand the basic fundamentals of disease.

What does caBIG do?


caBIG started in the US in an attempt to try to interconnect our NCI supported cancer centres. These are groups that are responsible for both conducting cutting edge research at a basic science level all the way through to delivering state of the art care. So we started this in an effort to try to do their research work but more recently we're moving forward to help them connect all their healthcare delivery information as well.

What do you think are the main things caBIG has to offer researchers – is it the collaborative aspect, the data management, working across borders..?

Part of the complexity of the problem we're trying to address is all the above..at its simplest form this grid or e-infrastructure gives people access to data and, or analytic capabilities that they just wouldn't have in the absence of being able to plug into this type of tool. On the other hand it also provides them with capabilities to manage their own data, to connect that data with other people's data to be able to then share their results and to form collaborative teams, to create virtual organisations. One of the key things that the founders of grid technology talk about all the time is that they enable the concept of virtual communities, of virtual organisations that supercede the technology, and even supercede the analytics. It's a whole new way of thinking about organisations that don't require you to have bricks and mortar or formal affiliations but in fact by using a technology platform and by using a series of trust agreements you can work together in ways that are just unprecedented.

What do you hope for the future of eHealth and caBIG?


Our hope is that we're going to change the face of biomedicine in general, but, in specific, move much more rapidly to being able to intervene to remove the tremendous burden of cancer. In the US alone there are 1.4/1.5 million diagnoses of cancer a year...Everybody is hysterical about H1N1 [swine flu] right now and the number of deaths is measured in low thousands at this point whereas in the US alone there's 500 000 deaths due to cancer. Our goal is to release the power in these next generation data sets, to liberate the information that's trapped in individual laboratories or in individual organisations so it can be more widely shared and so we can much more rapidly make progress in defeating cancer and all sorts of other diseases.

Wednesday, June 4, 2008

OGF23 podcast - Future challenges of grid computing?

Closing comments on the challenges facing grid computing in the future, starring Mario Campolargo (from the European Commission), Craig Lee (OGF president), Dieter Kranzlmueller (from the European Grid Initiative), Carlos Henrique de Brito Cruz (Scientific Director of the State of Sao Paulo Research Foundation), and Francesc Subiradia (Associate Director of the Barcelona Supercomputing Centre)...

This concluded the media conference...thanks to everyone involved!

OGF podcast - Is investment in grids paying off?

Two excellent answers -- from Mario Campolargo and Francesc Subiradia -- to a tricky question: Is investment in grid computing paying off?


OGF23 podcast - Mario Campolargo on the global inpact of grids

An excellent answer to a tricky question: Has investment in grid computing paid off?



OGF23 podcast - Craig Lee on grids for non-specialists

Will grid computing ever catch on in the public arena?

Craig Lee, president of the Open Grid Forum, speaks here on grids and clouds for non-experts...

OGF23 podcast - Mario Campolargo on grids in Europe

Mario Campolargo of the EC hit off yesterday's press conference with some words on grid computing in Europe...unfortunately, we had some sound problems during the recording, so this is just the audible bit of what he said...

The big word in his presentation was sharing: Grids enabling countries to share power, share data, share talent and people power... He also pointed out that grids are "one of the important vectors" in empowering the future of Europe.

OGF23 GridCast - Grid computing in Brazil

An inspiring podcast from Carlos Henrique de Brito Cruz, Scientific Director of the State of Sao Paulo Research Foundation, who speaks here at the OGF23 press conference about grid computing and eScience in Brazil.

Monday, June 2, 2008

Les Robertson kicks off with a keynote

A nice keynote from Les Robertson, who’s been heading up the LCG project, an effort to pull together a grid of 100,000 computers from institutions around the world in time for the start-up of the Large Hadron Collider, a particle accelerator that will create around 15 petabytes of data every year. [Shameless self promotion: check out this terrific interview with Les by me in iSGTW, complete with picture of Les from 1974 ;-)]

New to the LHC?

The LHC (Large Hadron Collider) is a 27-km ring of superconducting magnets that will generate temperatures a billion times hotter than the heat of the sun by smashing together particles traveling at 99.999999991% the speed of light (nice!). It’ll do this around 40 million times every second, which is more often than my Internet dropped out during this session. And it will operate at 1.9 Kelvin, which is minus-a-lot when converted to Celsius, and colder than temperatures in outer space. The LHC cost about 3 billion Euros and is designed to help us answer some big questions about our Universe.

The LHC Computing Grid (LCG)

Known to many as the lab “where the web was born,” CERN is now leading the LCG project, an effort to make the data generated by the LHC available to physicists around the world, in almost real time. Just one of the LHC experiments involves 2000 physicists from 34 countries and 150 universities.

Why choose grid for the LHC data?

1) each of the gazillions of particle collisions that will happen inside the LHC are independent events, which means you can analyze each event on an independent computer. This is a big “thumbs up” for grid computing.

2) the codes required to analyze the events are pretty small in terms of the memory they take up (around 2 gigabytes), which means you can use ordinary PCs to do the analysis. Another big “thumbs up”.

Plus, the nature of a computing grid is that it is distributed: scientists across the planet can access this data from the convenience of their own office, in their own lab.

How's the LHG progressing?

In Les’s own words: “So far, so good.” The LHG currently involves more than 140 computer centres around the world, including 60 federations in 35 countries. It averages around 300,000 jobs a day, which is the same as running 35,000 Intel cores 24 hours a day, seven days a week. (YAY for automation!).

And how important is this computer grid to the success of the LHC?

Essential. The distributed system must work from Day 1 of the LHC beginning collisions.

Does it work?

Les’s words again: “We’ll see when the physics comes out. We’re in the process of doing some final testing.” It’s all looking good at this stage: All of the baseline functionality is deployed and in use, so the focus of testing is now on performance and operational issues, like data distribution and access and storage management.

But, Robertson stressed that there are no guarantees. “The LCG is research. Until the real data comes we don’t actually know what people are going to do. We’re certainly looking forward to very exciting times.”

Yee ha! I quite like exciting times.