Pages

Monday, July 1, 2013

Asynchronous Parallel Computing Programming School in Bucaramanga, Colombia



With more than 70 students from different South America countries and the support of the Barcelona Supercomputing Center, Spain (BSC), the Asynchronous Computing Programming with MPI/OMPSs addressed to hybrid architectures school is developed in Bucaramanga, Colombia.  The school is organized by the High Performance and Scientific computing Laboratory of the Universidad Industrial de Santander (SC3 UIS) in Bucaramanga, Colombia, and continues until next Friday 5.

The school search to diffuse specific programming competences to the researchers, engineers and students which interact with the Latin-American and Caribbean Service of advanced computing (SCALAC from Spanish/Portuguese acronym), specifically with hybrid architectures as GUANE-1 the main HPC platform of the SC3 UIS (http:// sc3.uis.edu.co )
Several applications to test in this school, are related with particular uses in science and engineering, for example, weather applications, bioinformatics and computational chemistry, astrophysics, condensed matter, energy and seismic. SCALAC joint all Grid and Advanced Computing experience received in some projects developed in the last 10 years in Latinamerica.


More information: http://www.redclara.net/indico/evento/ompss 

Friday, June 28, 2013

iSGTW teams up with NUANCE to increase coverage of Africa

iSGTW is extremely pleased to announce that it has signed a memorandum of understanding with NUANCE, allowing the limited sharing of some content between the two publications.

NUANCE stands for ‘The Newsletter of the UbuntuNet Alliance: Networks, Collaboration, Education’ and is a publication we at iSGTW hold in high regard for its excellent coverage of national research and education networks (NRENs) in Africa.

At iSGTW, we hope that this exciting new partnership will allow us to increase our coverage of this region, where many exciting developments in the world of e-infrastructures are currently taking place.


You can read the latest edition of NUANCE on the UbuntuNet alliance website, here.

Wednesday, June 19, 2013

"Moore's Law is alive and well" — but is that enough?

On Monday, with the announcement of the new Top 500 list of the world's fastest supercomputers, we wrote briefly about the challenges computer scientists across the globe face in achieving exascale supercomputers by the end of the decade. To put the scale of this challenge into perspective, China's Milky Way 2 supercomputer, the fastest in the world today by a significant margin, is capable of reaching 34 petaFLOPS. Plus, there's the small matter of energy efficiency still to tackle if exascale supercomputers are going to become a realistic proposition.

Yesterday evening, Stephen S. Pawlowski of Intel gave a keynote speech at ISC'13 entitled 'Moore's Law 2020'. "People are always saying that Moore's Law is coming to an end, but transistor dimensions will continue to scale two times every two years and improve performance, reduce power and reduce cost per transistor," he says. "Moore's Law is alive and well."

"But getting to Exascale by 2020 requires a performance improvement of two times every year," Pawlowski explains. "Key innovations were needed to keep us on track in the past: many core, wide vectors, low power cores, etc."

"Going forward, scaling will be as much about material and structure innovation as dimension scaling". He cites potential future technologies, such as graphene, 3D chip stacking, nanowires, and photonics, as ways of achieving this.

Pawlowski argues for less focus on achieving a good score on the Top 500 list by optimising performance for the Linpack benchmark. Instead, he says, there needs to be more focus on creating machines suited to running scientific applications. "Moore's Law continues, but the formula for success is changing," concludes Pawlowski.

Monday, June 17, 2013

Top of the FLOPS at ISC’13

This week, almost 2,500 experts from industry, research, and academia have gathered in the German city of Leipzig for International Supercomputing Conference ’13 (ISC’13). The event played host today to the announcement of the new TOP500 list of the fastest supercomputers in the world. Milky Way 2 (known also as Tianhe-2), located at the National University of Defense Technology (NUDT) in Changsha, China, was announced the new winner. “The Milky Way 2 project lasted three years and required the work of more than 400 team members,” says Kai Lu, vice dean of the School of Computer Science at NUDT. Boasting over 3 million cores and with a peak performance of around 34 petaFLOPS on the Linpack benchmark, Milky Way 2 is nearly twice as fast as the previous winning supercomputer, Titan, at Oakridge National Laboratory, US. Titan has now slipped to number two spot on the list, with another US-based supercomputer, Sequoia, located at Lawrence Livermore National Labs, completing the top three. JuQUEEEN at the Jülich Supercomputing Centre in Germany was ranked as the fastest machine in Europe.

“Our projections still point towards reaching exascale systems by around 2019,” says Erich Strohmaier of the US Department of Energy’s Lawrence Berkley National Laboratory, who gave an overview of the highlights of the new Top 500 list. Strohmaier, however, warns that increasing the power efficiency of supercomputing systems will continue to be a major challenge over the coming years: “If we don’t start to have some new ideas about how to build supercomputers, we will truly be in trouble by the end of the decade.”


 “If you actually look at what people want to do, an exaflop is still not enough,” says Bill Dally of NVIDIA and Stanford University, California, US.  He capped off this morning’s programme with a keynote speech on the future challenges of large scale computing. “The appetite for performance is insatiable,” he says, citing work in a number of research fields as evidence that performance is currently still the limiting factor in terms of the exciting science which can potentially be done. “If we provide increased performance, people will always find interesting things to do with it.”

Latinamerican High Performance and Grid Computing Community calls for contributions to CLCAR 2013 in San José Costa Rica

The Latinamerican Conference on High Performance Computing (CLCAR, from spanish acronym) 2013 will be held this year in San José, Costa Rica. Since 2007, the Latin-American Conference on High Performance Computing (CLCAR) is an event for students, scientists and researchers in the areas of high performance computing, high throughput computing, parallel and distributed systems, e-science and applications, in a global context, but with special scope in latinoamerican propositions. GridCast is media-partner of this latinamerican activity.

The program and scientific committees are formed by experts and researchers from different countries and related domains. Competent people from various countries and institutes will carry out the process of evaluating the proposals. CLCAR 2013 to be held in San José,  Costa Rica, in August 26-30. 

CLCAR 2013 official languages are English, Portuguese and Spanish. People willing to present their proposals can present them mainly in two forms: Oral Presentations (Full Paper) and Posters (Extended abstract) until Sunday, June 23.

This year there are two activities proposed inside the conference: the first one, the bioinformatic and biochemistry researchers propose the bio-CLCAR, and the second, the CLCAR scientific visualization challenge. 

Different Proposals can be  submitted in  ENGLISH, PORTUGUESE or SPANISH only (full papers and extended abstracts).  Papers written in Spanish or Portuguese should have the title and its abstract in english too. The oral presentation may be in any CLCAR official language, but the slides will be in English anyway.  Selected posters from extended abstracts must be show in English.

For more information about CLCAR 2013, please visit the official site: www.clcar.org

Praise for PRACE and the importance of building expertise in HPC


Yesterday, the PRACE Scientific Conference was held in Leipzig, Germany. It is one of several satellite events taking place alongside ISC'13, which gets underway in full today.

After a brief welcome address from Kenneth Ruud, chairman of the PRACE Scientific Steering Committee, Kostas Glinos, head of the European Commission's eInfrastructures unit, spoke about the vision for HPC in Horizon 2020.

"HPC has a fundamental role in driving innovation, leading to societal impact through better solutions for societal challenges and increased industrial competitiveness," says Glinos. "It's not just about exascale hardware and systems, but about the computer science needed to have a new generation of ICT."

"Only very few applications using HPC really take advantage of current petaFLOPS systems," he adds. "New computational methods and algorithms must be developed, and new applications must be reprogrammed in radically new ways." In addition, Glinos  highlighted the importance of public procurement of commercial systems for developing the next generation of IT infrastructures, which you can read more about in the recent iSGTW article ‘Golden opportunities for e-infrastructures at the EGI Community Forum’.

Finally, he spoke about the conclusions of the recent EU council for competitiveness: "HPC is an important asset for the EU... and the council acknowledges the very good achievements of PRACE over the years." For Horizon 2020, Glinos says: "We want to build on PRACE's achievements to advance further integration and sustainability." He argues for the importance of an EU-level policy in HPC addressing the entire HPC ecosystem, saying that the sum of national efforts is not enough – "we need to exchange and share priorities."

The conclusions of the EU council for competitiveness were also highlighted by Sergi Girona, chair of the PRACE board of directors. "We have to work together because we want to support science and industry, the development of HPC in Europe, and the development and training of persons," he says.

During his talk, Girona also gave an overview of PRACE in numbers: with its 25 member countries, PRACE has a budget of €530m for 2010-2015, including €70m of funding from the European Union. Girona explains that PRACE has now awarded more than 5 billion computation hours since 2010 and is currently providing resources of nearly 15 petaflops.

However, he emphasises that PRACE is about much more than simply providing access to HPC resources. "We don't just want to give access to computing resources; we want to support users at all stages – it is key to train people," he says. "We have created six training centres in Europe and have approved a curriculum with 71 PRACE advanced training centre courses for this year."

The importance of training was also highlighted by Glinos: "We need more expertise, so we intend to support a limited number of centres of excellence. Topics may relate to scientific or industrial domains, such as climate modelling or cancer research for example, or they may be 'horizontal', addressing wider challenges which exist in HPC. These centres of excellence need to be led by the needs of the users and the application owners."

Following Girona's talk, Wolfgang Eckhart of the Technical University of Munich, Germany, gave a presentation on his research in the field of molecular dynamics. He and his colleagues have been selected as winners of the PRACE ISC Award for their paper entitled '591 TFLOPS Multi-Trillion Particles Simulation on SuperMUC'. The award ceremony is set to take place later today.

The remainder of the conference consisted of a series of exciting presentations on research conducted using PRACE resources, ranging from high-resolution global climate models to molecular simulation, and from astrophysics to better understanding the building blocks of matter. You can read more about these on the PRACE Scientific Conference website, here.


Be sure to check back later this week for further updates from ISC'13.

Tuesday, June 11, 2013

IT as a Utility in Emerging Economies

Mobile is critical for IT in emerging economies.
(CC-BY-NC-SA AdamCohn, Flickr)
The ITaaUN+[1] workshop on IT as an infrastructure in emerging economies attracted social activists-cum-academics, academics-cum-industrial consultants, linguists, digital humanitists, and technology visionaries to the Association of Commonwealth Universities (ACU), where ACU Secretary General, John Wood, played host to the compact but vocal group. The agenda was to discuss the challenges and opportunities of IT, seen as a utility, in the majority world. John Wood is himself a veteran of e-infrastructures in the UK, having being Chief Executive of the Council for the Central Laboratory of the Research Councils, where he was responsible for RAL and Daresbury. Later, he held positions at Imperial – first as Principal of the Engineering Faculty, and later as Senior International Advisor. He now sits on the board of JISC and the British Library, and has advised numerous governmental and corporate organisations across the globe. But his experience at the ACU gives him a unique perspective on infrastructures that are in place already in the Commonwealth countries that are also developing countries (assuming provision of computational infrastructure in higher education institutions is an accurate barometer of infrastructure elsewhere in countries, which it usually is, to some degree).

Why the service/utility distinction though?

Jeremy Frey, Physical Chemist at the University of Southampton and one of the minds behind Chemistry2.0 application CombeChem, explained that there is a natural progression of a technology as it becomes part of the fabric of our lives. The transition: Revelation > Innovation > Specialist Tool/resource > Service > Utility – is one that the utilities that form the infrastructure of daily activity, such as electricity and telecommunications, have already progressed along. IT, and especially networked IT, is somewhere between the last two stages (and there will continue to be debates about whether and to what extent all utilities should be services or utilities, depending on the economic model in place). But, as the UN’s World Summit on Information Society (WSIS) has suggested, internet access is beginning to be established by governments as a basic human right, and so it will become increasingly considered a utility rather than a service – something we need rather than just something we want. And the impact that this will have will be felt nowhere more so than in the emerging economies of the developing world.

In 1965 the centre of gravity of global wealth was located on the plains of La Mancha, in Spain. This was, however, the last time that this topographical El Dorado[2] lay in Europe. As the years have passed, the centre of international wealth has meandered out across the Mediterranean, and zig-zagged through northern Algeria and Tunisia. Having bounced off Malta and back towards Tunis, it is now poised to zoom eastward once more, skirting past northern Libya to settle in the middle East some time around the middle of this century (remaining oil reserves notwithstanding).

The economic and social conditions that led to Europe being the dominant force in the world over the last few centuries can be attributed to a combination of factors: a great abundance of natural resources – particularly wood, coal and iron – suited to the manufacture of ever-more-complex tools; the concentration of historically competitive cultures within Europe itself (leading to wars, but also leading to a wealth of art, science and advanced technology); and the suitability for domestication and farming of indigenous fauna and flora. This last blessing, in combination with the large Eurasian landmass whose principal axis is East-West and therefore at the same latitude, allowed successful farming methods to spread very easily over a large area[3], dramatically increasing population size in Europe and therefore the means to successfully colonise new areas of the globe.

With war and recession, and with growing social justice – helped in part by the ease with which injustices can be brought to the World’s attention through social media – the old colonialism has crumbled and many of its worst crimes are behind us. However, some corporations exploit the cheap labour and (often) rare natural materials of developing nations in much the same way as colonial powers once did[4]. Sometimes this is for food: growing soya for cattlefeed, for example…but potentially at the expense of rainforest. Sometimes this is for palm oil: a food and a fuel, but at the expense of rainforest and peat bogs, potentially leading to local and global environmental problems. And in many African nations, landscapes are plundered for rare metals, destined for today’s mobile devices.

Today’s mobile devices are not, however, solely owned by rich people in rich countries. Competition between mobile device manufacturers and an increased acceptance of their ubiquity in our everyday lives has led to dramatic price drops in the developed world. And together with greater prosperity in the developing world, such devices are now becoming more affordable for a growing global middle class – mainly urban; often involved in the knowledge economy in some way. People in these countries are leapfrogging the PC/browser paradigm, and (for a large number) the first personal device they own that can access the web and email is in fact a webphone or smartphone.

Although there is still tremendous disparity between rich and poor in the world, developing countries are experiencing this growth of a middle class (defined as having $10-$100 of disposable income per capita per day in PPP terms), exactly at a time when those people are being empowered with devices that help them to access information and to make decisions. The pace of change is rapid. New technologies are initiating and catalysing social change (in the manner of the Arab spring), and opening up new possibilities in the realms of education and research, communications, and e-governance. In many ways, ever-cheaper technologies are helping to correct the economic unevenness that led Europe (and North America) to become unduly dominant over the last few centuries. Mobile web as a globally democratising tool, if you will.

For that reason, the potential for real change in emerging economies that might be realised by IT has attracted visionaries and projects from the developed world, eager to demonstrate the power and potential of technology to enact measurable social change quickly. E-Governance, for example, is an area where both the potential for positive change and the motivation to do so, especially in countries that are newly democratic, is very attractive. But at the ITaaUN+ workshop, individuals with direct experience of having participated in such programs suggested that this approach can seem like neo-colonialism. Why implement (or as it might seem, impose) e-voting tools, for example, when similar innovations haven’t always had brilliant success in long-established democracies? It could smack of using developing countries as test cases for technologies that haven’t already been tried and tested in the developed world.

There are, however, good examples of IT making positive changes to people’s everyday lives. Take microcredit, for instance. This allows people to send small amounts of currency using their mobile phones. By using SMS as the technology behind the system, microcredit uses a simple, tried-and-tested service that, though perhaps lacking the bells and whistles of wifi, is more reliable in sub-Saharan Africa, where internet connections cannot always be guaranteed. Distributed systems, whether for power generation through smaller, local solar and wind power; or computer networks themselves – might offer solutions that bring specific benefits to emerging economies, where geographical isolation can be a barrier to some technologies being more widely adopted. 

The discussion next led on to 3D printing. The 3D4D challenge, for example, looks at the benefits this rapidly developing technology could bring to emerging economies. From simple to complex devices (including medical tools), to an enabler of innovation, 3D printing could have a huge impact on emerging economies, although much interest has come so far from the developed world. The ability to make new parts and eventually more complex components could mean that emerging economies, which are already embracing a more sustainable approach to technology adoption through frugal innovation, manage to avoid some of the more wasteful consumer trappings of buying every new model that are prevalent in the West. As resources needed to make new mobile devices are depleted, the developed world could learn a thing or two from the emerging economies.

The link again for ITaaUN+ is: www.itutility.ac.uk

...they're running workshops and events all the time, so it's worth taking a look.


[1] Information Technology (ok, you probably knew that bit) as a Utility Network-plus
[2] There are several maps tracking the movement of the centre of economic gravity, over a number of time-scales. I’m using one by Homi Kharas of the OECD development centre: working paper #285, “The Emerging Middle Class in Developing Countries”. Although it simplifies the calculation, by assuming the centre of each nation’s GDP is located on its capital city, this is a fairly reasonable assumption to make (better, in my opinion, than assuming that GDP is colocated with the geographic centre of each country). There are, of course, countries where the capital is not the wealthiest city – but in many cases these are close enough geographically (e.g. New York and Washington; Toronto and Ottawa) to not make too much difference. Sydney and Canberra are ‘close enough’ at the distance they are  from the centre; Milan–Rome and Köln-Düsseldorf–Berlin present more of a problem due to their proximity to the centre, but this is still the best solution I have seen.
[3] 1. This argument is from Jared Diamond’s Guns, Germs and Steel
2. Crops are generally sensitive to being north/south variance but will grow happily along the same latitude
[4] What corporations are not able to do is maintain control over entire nations, so that, provided governments in developing countries offer some degree of freedom to their citizens, individuals and groups are able to act in an entrepreneurial manner

Wednesday, June 5, 2013

Jazz music and Big Data at the TERENA Networking Conference 2013

Nearly 20 years ago, the European Union came into being – this week’s TERENA Networking Conference is taking place in Maastricht where the treaty that created it was signed. Regardless of whether you see the EU as a positive or negative entity in today’s cash strapped times, it’s appropriate that the Trans-European Research and Education Networking Association should meet here to discuss the next wave of Telco innovation. Hopefully this will mean a few extra Euros will be on their way to all of us in the future.

The Opening Reception on Monday featured a jazz performance inspired by the theme of the conference - "Innovating Together". Thanks to an international collaboration of artists and technicians, on-stage musicians performed alongside their 'virtual' band leader, who was in Edinburgh, UK, assisted by LoLA (LOw LAtency audio/visual system). LoLA was developed by GARR, the Italian research and education networking organisation, and the Music Conservatory G. Tartini in Trieste. Using LoLA, performing artists are able to interact in a natural way even if they are thousands of kilometres apart, relying on the high-quality and very large bandwidth connectivity offered by research and education networks which minimise network-related delay and jitter.



 One source of innovation is likely to be data, and lots of it. In a session called ‘Big Data, Big Deal!’, Harold Teunissen of SURFNet looked at the big data problem. He noted that after the arrival of his twins, he found himself faced by a big data problem of his own – over 30,000 family snaps to share, store and manage. “These two changed my big data perspective for ever,” he said.

In the Netherlands, all ICT activities for Higher Education and Research are now brought under the SURF umbrella including cloud, supercomputing and the esciencecentre. But what is big data? NRENs generate 0.3% of worldwide data. Is this data actually big data, or is mainly people using Facebook, Twitter etc. We don’t know. Research is now seen as a generator of big data, but the cost of generating it can be very high. For example, Teunissen noted that 10 billion dollars spent on the Large Hadron Collider to arrive at one bit of information might be seen as a lot of money by some i.e. proving that Higgs particle exists = true. Of course, the story of the LHC and its research is a lot more complex than that.

Big data actually means large volume, generated at a high velocity and in many different forms, such as video, text and images. What customers need to handle this data tends to either be technology and performance OR solutions and ease of use, depending on whether or not they are early adopters.

Taking up the LHC theme, William Johnston of ESnet looked at high energy physics as a prototype for data intensive science, now paying dividends for the teams working on the  Square Kilometre Array and genomics for example. Growth in scientific data follows Moores Law, leading to exponential growth. However, when looking for technological solutions to big data problems, commercial solutions may not be up to the task.“Software testing started 5 years before the LHC turned on. Science is not YouTube and has special requirements,” said Johnston.

Simon Leinen of Swiss NREN, SWITCH asked whether we should in fact make science more like YouTube, so explore using the cloud and existing services? Johnston responded that HEP is looking at its requirements and how these map onto commercial services, but highly parallel flows of data are somewhat unique to HEP (You can read an article about CERN’s cloud choices by David Meyer in Gigaom at http://gigaom.com/2013/05/31/heres-why-cern-ditched-opennebula-for-openstack/). "This may not be the case for other fields, such as environmental modelling", pointed out Johnston. "Joining supercomputers together to work on a problem is not the same as parallel computing."

Tuesday, May 28, 2013

Mapping ICT accross Sub-Saharan Africa: iMentors


What level of connectivity, types of networks, data infrastructures, e-tools, and projects etc.are currently available for researchers  in Sub-Saharan Africa?  One European FP-7 funded project is starting to map this knowledge in a virtual observatory. 

The idea behind the i-Mentors project was first conceived by the Department of Computer and Systems Sciences at Stockholm University and Gov2u.  Although Sub-Saharan Africa has witnessed dramatic growth in Information Communications Technology (ICT) access since the mid-1990s, there are still gaps in the knowledge around the status and findings from past and ongoing e-infrastructure projects. iMentors now plans to map the entire e-infrastructure landscape accross Sub-Saharan Africa recording key developments over the past five years. Their goal is to provide valuable insights to policy makers, international investors, educators and researchers. Any gaps and/or progress will be recorded and collated in order to enhance ICT initiatives in the region.

Our project, e-ScienceTalk, is delighted to have just signed an Memorandum of Understanding with iMentors. We are hoping to help increase the visibility of the project's findings and success stories. Check out their Facebook page here and Twitter stream (@i-mentors). 


The project's first step is to gather valuable stakeholder feedback to define the most appropriate criteria in evaluating e-infrastructure projects. So please do help them by filling in their survey (http://bit.ly/YF5GBV).
 

Friday, May 24, 2013

Núria Bel at e-IRG May 2013 Dublin

Núria Bel works at the Institute of Applied Linguistics at the Universitat Pompeu Fabra in Barcelona. She talked to us about e-infrastructures for linguistics at the e-IRG workshop in Dublin.


Sverker Holmgren on the e-IRG workshop in Dublin

e-IRG Chair Sverker Holmgren had some time to talk to us about the workshop in Dublin. Here's what he had to say:


Thursday, May 23, 2013

Donald Knuth said software is hard. So is open data.

Software is data. Data is infrastructure. But is data research, or is it development? Is it the foundation for an academic career, or is it a raw material, ready to be processed and commercialised by the entrepreneurs of the digital age?

If data is the new oil, then we need to know who can lay a credible claim to it. Taxpayer-funded public research, it is widely said, should be made openly accessible. But researchers want credit for their hard work and, concerned about their data-well being plundered by unscrupulous others, want to control access to it.

This database might be copyrightable. And quite large.
CC-BY-NC ~BC
Raw data cannot be copyrighted: it fails to meet the criteria of being creative and (in its representation of an aspect of the physical laws of the Universe) original. Databases, in their curation and design, are protected by thin copyright – and so the intellectual property owner can decide how the database can be reused. That means they have copyright, but can choose to licence it under copyleft if they so choose, e.g. Creative Commons attribution or similar.

I’ve deliberately omitted specific pieces of information in the last paragraph. First: who is the intellectual property owner? (Back to this point shortly) Second: why bother licensing data that, in the true gold standard for open date, should receive a public domain dedication or, in the Creative Commons world, a CC0 licence? That’s what the Panton Principles recommend: making data truly public domain avoids the unsure nature of thin copyright (by simply renouncing ownership) and also avoids the eternal devil of Creative Commons attribution works… attribution stacking (if the data is produced by seven authors and is then recycled six times by author-groups of seven, all different each time, it will effectively have 49 authors, and this can grow).

Depending on the country and institution, it may be that the University (or research institute) is the real copyright holder (if the database is curated and copyrightable) rather than the researcher. That’s surprising even to some researchers who, in signing a contract, relinquished intellectual property rights when they took a job at University X. But clearly, researchers can be copyright holders and therefore (and especially if they are funded by taxpayers) the common good requires them to make their data open, i.e., PDDL or CC0. After all, their hard work and the credit they deserve from it both derives from and is protected by cultural norms: citations, election to learned societies and so on. But, as I learned at the e-IRG workshop in Dublin, not all publicly funded scientists do this: they are worried that their peers will plunder their data-well and not cite the source, either absent-mindedly or (in the worst case), maliciously. So they release their original, creative database under CC-BY, or something more restrictive, in direct contradiction to the Panton principles. Or they don’t, and presumably assert copyright (whether their work is copyrightable or not).

Research institutes can also be as confused in their approach to licensing, and many legal experts in the field recount examples of institutions worrying about licensing after the fact (when it is often too late). But, for simplicity’s sake, those advocating open data would rather the intellectual property resulting from publicly funded research be in the hands of institutes rather than individual scientists, when it is easier to manage.

In the open data legalities track of the e-IRG workshop, it was suggested that a lot of confusion can be avoided by better training in such matters as copyright and copyleft for researchers, and default licensing positions for publicly funded data. Having a default position recognises that many researchers shouldn’t have to care about legalities if they don’t want to; better education about copyright and copyleft, I believe, would make everyone better web citizens. Researchers whose databases are used deserve credit, but proper citation of source also deserves credit if we are to ever move towards public domain dedication. Databases used without proper citation in research to further a career, when carried out duplicitously, is one thing only: plagiarism. That alone should dissuade unscrupulous researchers, if they really exist.

The onus is also on the intellectual property owner to flag up instances of data misattribution if it occurs, but we should work towards a persistant identifier system, so you can track the data back to source.

So: that’s publicly funded research. Now, what about the private sector? Or, indeed, the start-up that the researcher runs in their spare time: who should own that? The individual? The University, if the individual thinks about the start-up between 9 and 5 on a weekday? It gets tricky. But better education in these matters, and actually having a default position, is a step towards a more sensible future with regard to open data.

Oh, and regarding software: not copyrightable. But licensable. That's a whole other argument.

Tuesday, April 30, 2013

A new era for the post EMI: all together for MeDIA

The European Middleware Initiative (EMI) presented a new initiative for a long-term, open, lightweight collaboration on the coordination of distributed middleware technologies: the Middleware Development and Innovation Alliance (MeDIA). The goal of the initiative is to facilitate the future development and evolution of middleware solutions beyond the current short-term project limits. The event launch took place last week in one of the most ancient places of the world, Rome, just remembering what Roman poet Ovid said in his Metamorphoses: omnia mutatur; nihil interit, that is everything changes, nothing perishes. The place was indeed the right stepping stone to open a new era for the EMI project, which achieved many important results in the design and implementation of common services and technical agreements in several areas.
As EMI has shown the importance of coordination and iterative practical implementation and validation, MeDIA would like to provide this same highly qualified cooperation mechanism through the definition of dedicated working groups based on contributions from several development teams. The workshop represented a truly open kick-start meeting to summarise three years of work and achievements of the EMI project and to discuss about future goals, scope, activities finalised to setting up new technical collaborations among team leaders, members of middleware development teams and other interested parties, such as technical experts from the user communities, infrastructure providers, application developers, and commercial companies.
During the entire event, the heated discussions put the spotlight on an effective and sustainable continuation of the key activities. In substance, MeDIA plans to provide a forum for coordinating the innovation and development of middleware services across academic and scientific research infrastructures based on the members’ interests and priorities. The members’ participation is voluntary and bottom-up and relies on an active sharing of information, proposals and ideas from members to members.
The participants were very proactive in all debates and, after fruitful exchanges, have outlined the following goals for the MeDIA initiative:

·         To provide a forum for synchronization of technical development roadmaps;

·         To concretely act upon the actions identified during the roadmap discussions;

·         To build the bases for a solid collaboration platform on distributed computing and data management;

·         To enhance the relationships among all partners using modern social networking techniques.

In conclusion, MeDIA is intended to be not only a coordination initiative to give software developers from existing EMI product teams, but also a place where the technical collaboration can continue and expand to include other development teams worldwide. Last, but not least, this initiative aims to outline future technical roadmaps for the European research infrastructures, promote interoperability and common development work, and put in place all the necessary activities to continue the success of EMI.

Friday, April 19, 2013

Wrapping up at CAPRI – beers, bikes and brainstorming

At the end of a long day at the CAPRI meeting, the sessions closed with a brainstorming discussion on the present and future of CAPRI. Which targets should be in the frame for the next phase? Before the discussions got underway though there were a series of lightning talks from early stage researchers – great to see them brought into the spotlight. A session on docking methodology and servers followed. Paul Bates of Cancer Research UK focused on the SwarmDock server, a webservice for generating 3D structures of protein-protein complexes, by looking for the lowest energy solutions. He described the swarming process as a bit like wandering around Utrecht after a few beers – you might end up in the canal, get lost down a cobbled street or “be run over by one of these high speed cyclists.” There was a big laugh of recognition from the delegates for that one.

Anisah Ghoorah from the University of Lorraine, INRIA, pointed out that, “Docking is difficult, but that also makes it interesting.”  Ab initio techniques i.e. starting from scratch are too difficult, so there’s a need to use templates and knowledge based approaches. “CAPRI should focus on the difficult cases – leave Google for the simple ones,” he said.

ClusPro was the first fully automated, web based program for computational docking of protein structures. It has been around for 10 years, and has about 3000 users. Sandor Vajda of Boston University had been looking at usage of the server over the course of the year.  “First of all you can see that dockers don’t work on the weekend – and they go on holiday in August,” he pointed out. “And no one cares about Christmas!”

Martin Zacharias of TU Munchen wound up the session by presenting protein-protein docking and refinement with a coarse-grained protein model. The computing needed seems very efficient. “It’s possible to do 10,000 docking simulations on a single core in an hour,” said Zacharias. “These happen much faster in reality than they do in silico.”

Finally, the community took at look at some new areas proposed for CAPRI. These include interface predictions, useful for cancer studies and interface water predictions, which are potentially very challenging and are currently underrepresented. Also floated were protein-peptide interactions, protein-oligosaccharides, protein-RNA and even protein-DNA (with a question mark). For affinity challenges, it’s a question of what they both can and want to evaluate. From the discussion, it was clear that they are looking for a mix of easy and difficult targets, as well as having lots of targets to choose from. However, cooperation from the structural community is essential – they need to submit their structures to compare with the predictions. “They won’t solve the targets just to be nice to us!” it was pointed out. So it’s not just protein-protein interactions that are important in the end, but those between the structural biologists and modellers as well.

Proteins by design – expanding on nature

Yesterday at the CAPRI meeting,  a heartfelt plea from Dr Ilya Vakser of the University of Kansas really caught my ear. “There’s so much data and so few people to dig into it!” he cried.

Dr David Baker of Washington University, Seattle introduced the CAPRI meeting to the design of proteins, and how they interact with each other. It could all have been different though – first Baker polled the audience to see whether they would prefer to hear about structural modelling with sparse data instead.  That’s the first time I’ve ever seen audience participation in a talk at that level – it was practically the conference equivalent of X Factor.

Once he had the correct set of slides cued up, Baker started out by praising the excellent job that Mother Nature’s naturally occurring proteins do. They solve the challenges that biological evolution presents perfectly, such as capturing and storing energy, making and breaking down molecules. However, thanks to modern life we now find ourselves faced with challenges that are not thrown at us by nature alone, such as global warming and living longer. The question for protein designers is can they design a whole new world of synthetic proteins to address these challenges? “Ideally we should take less than 2 billion years to figure out how to do it,” said Baker.

There are a number of areas to focus on, including designing the next generation of therapeutic proteins, building scaffolds for enzyme reactions, and self assembling cages to carry drugs and vaccines into the body. The process starts by calculating a sequence of biological building blocks known as amino acids that might give the desired protein structure and function. You read off the sequence, and translate that back into a DNA sequence that would encode the synthetic protein. Once you have a DNA sequence, you can then make the gene that will then build the protein you want. As a non-biologist, this was the part that I found like an extract from a sci-fi novel – you can now buy a custom gene over the internet, and then back comes the DNA encoder that you need. Build your protein, then see if it does what you want.

But the process starts with finding the sequence with the lowest energy configuration – because molecules like to exist at the lowest possible energy. It’s difficult to know whether the structure you have found is the lowest possible however and this is where non-scientists can get involved. Rosetta@home is a volunteer computing project which uses spare computer time donated by volunteers to look for the lowest energy versions of protein sequences and send back the solutions. Luckily the synthetic protein structures are somewhat easier to predict than naturally evolved proteins.

Baker ran through some more key challenges for protein builders, such as designing binding between proteins when looking for ways to treat Spanish flu, and finding methods to inhibit intracellular interactions when you need to target just one black sheep out of a whole family of proteins, for example the Epstein-Barr virus. 

FoldIt puzzle screenshot
One question Baker asked is, "Why are the enzymes we design ourselves so lousy?" Here he turned to the FoldIt volunteer community. FoldIt was dreamed up in response to Rossetta@home users who wanted a bit more of a challenge. FoldIt poses protein folding puzzles to users, to check protein design sequences and see if in silico folding predications are accurate.

One example is the de novo designed Diels-Alderase, an enzyme for forming carbon-carbon bonds.  "Can we improve activity by remodelling active site loops?" asked Baker. "Let’s ask the Foldit community!" He gave them a solve it game for FoldIt players and one successful player increased the interaction 20 fold. When the crystal structure was solved, it was found to be very similar to the prediction. ”Now I’ve come to be a big believer in the concept that if you don’t know the answer, ask more people until you do!” said Baker. A response to Vaksar’s plea earlier perhaps – the people to analyse your data are out there, you just need to give them the tools to do it.


(For more information about volunteer computing, read the e-ScienceBriefing at http://www.e-sciencetalk.org/briefings/19/ESTB-19-Desktop-grids-Connecting-everyone-to-science.html)

Combining and conquering at the CAPRI 5th evaluation meeting

After the beer and curry-fest of Manchester and the EGI Community Forum last week, e-ScienceTalk is taking a more leisurely sojourn in a sunny, cobbled Utrecht for the CAPRI 5th Evaluation Meeting. The strapline for the event is ‘Combine and Conquer'. Many processes in the body’s cells are driven by large, complex molecular networks of proteins. The 3D structures of these combined complexes are often the key to their function. But deciphering their detailed structure is difficult, because you need access to large pieces of kit such as X-ray synchrotrons or  nuclear magnetic resonance machines, plus these protein complexes can be fiendishly tricky to crystallise. Lower resolution techniques such as small angle X-ray scattering can give clues and increasingly this data is being used as a starting point for computational modelling. Predicting, modelling and understanding these large complexes is making an important  contribution to drug and protein design – but how confident can you be that the models are correct? CAPRI (Critical Assessment of Predicted Interaction) is an international effort to assess the performance of these methods, by inviting developers to test their algorithms on the same protein targets. This week’s meeting focuses on assessing the performance of docking methods in predicting the 3D structure of complexes and is extending the challenge to predicting binding and de novo interface design.

The opening keynote was by Piet Gros of Utrecht University and winner of the prestigious Dutch Spinoza research prize in 2010. Actually, Gros comes from Dokkum in North Holland, a neat coincidence with the focus on ‘docking’ interactions between proteins at the event. “Although if you think Utrecht is flat, it’s a hill compared to Dokkum,” he remarked.

Gros’s talk highlighted the complement system, which is an ancient part of our immune defence found in the blood. The complement system is formed by around 30 large plasma proteins and cell surface receptors. This system then recognises and eliminates bacteria, viruses and altered host cells, while protecting healthy cells. It links our inbuilt and acquired immunity. “Essentially, it cleans the garbage,” said Gros.

You won’t see a ‘complementology’ department in a hospital, but disruptions in this systems can lead to a range of health problems. It is linked to auto immune conditions such as rheumatoid arthritis, stroke and heart attacks and injuries due to a loss in blood supply, for example to the brain. Genetic changes in complement proteins can lead to kidney conditions and eye problems, such as age related macular degeneration (AMD).  They even have a role in infections, such as EHEC a type of E.coli bacteria, which led to the most expensive drug on the planet being used 2 years ago in Germany during an outbreak.

Using structural studies, Gros’s team has revealed the molecular mechanisms responsible for their functions, such as central amplification, protection of the host body and what happens at the start when a membrane-attack complex forms. In molecular terms, the structural re-arrangements involved are large – about 100 A when a typical atom is a few Angstroms in size (one angstrom is one ten billionth of a metre).

Alexandre Bonvin of the WeNMR project asked Gros to tell the computational modellers in the room what might be lacking from the toolkit he needs for his work. Gros didn’t pinpoint an example, but agreed that the docking modelling and structural work with X-ray crystallography were indeed complementary fields (no pun intended I guess). “In the end we should be able to model everything,” he said, “because otherwise it means we still don’t understand it.”

Over the classic local 'borrel' later, I asked Dr Gros about the computing he uses to support his structural work. Essentially he uses an in-house computing cluster, and does not foresee a need for expanding out to grid computing or the cloud at the moment. This makes an interesting contrast with the WeNMR project community, which relies heavily on grid portals to help them to work on NMR structural data. This presents a challenge for e-Infrastructure providers such as EGI.eu when reaching out to the ‘long tail’ of science – how to engage users with international federated resources when increases in computing power make home grown resources so attractively simple, but maybe not scalable? Does this mean that the long tail has no need of scalable solutions to solve their questions – or can the questions themselves scale up?

Tuesday, April 16, 2013

SCI-BUS update brings clouds to scientific gateways

On Friday morning, we had the chance to catch up with Zoltan Farkas from SCI-BUS and Pamela Greenwell from the University of Westminster, UK, to talk to them about the new development in SCI-BUS which is bringing clouds to the platform – and the new science this enables.


Friday, April 12, 2013

Tom Morrison from STFC: EGI Training Marketplace

Tom Morrison from STFC is the developer of EGI Training Marketplace. Here he shows us a demo of how it works:




Tom also had time to talk to us about his highlights of the Community Forum:

Thomas Kulhanek: EGI Champion!

Tomas Kulhanek talks to us about integrative physiology and his role as an EGI Champion:


Getting citizen scientists on your team

There are currently four million volunteer computers searching for new pulsars for the Einstein@Home. This sleepy distributed team has so far collectively processed over 1 petaflops. 46 individual pulsars have been discovered this way. Dr Ad Emmen (AlmereGrid) introduced a session yesterday, 'Getting citizens scientist on your team'. Ad roughly calculated that in Greater Manchester alone there could be over 2.5 million computer sitting idly. It is now estimated that there could up anyway up to 2 billion computers worldwide.  The potential growth of this computing capacity is enormous.

One project, International Desktop Grid Federation (IDGF), has been harnessing this processing power.  IDGF brings together those interested in desktop grids including 50 different desktop grids, 50 member orgnisations, and over 240 individual members. 

As an IDGF member, institutes have access to high level science gateways with a host of applications for end users. There is also certification support, a roadmap for guiding the management of desktop grids for administrators, plus a vimeo training channel and a technical wiki.

In June 2012, the monthly performance of the desktop grid virtual organisation (VO) in the EGI portal was an impressive 1,051,051 CPU hours.  It basically held the 10th spot for a while. However, this number fluctuates and IDGF have ambitions to capture more processing power for scientific research. By contrast, the ATLAS experiment uses 100,000,000 CPU hours.

Dr Robert Lovas from MTA SZTAKI introduced IDGF-SP . The idea behind this support programme is to gather experiences, promote success stories and set up a coordinated campaign to boost uptake of desktop grids by institutes, universities and researchers across the globe. It also hopes to encourage the growth of a network of citizen scientists.


What IDGF-SP is currently looking for are ambassadors to bridge the gap between scientists and citizens. Similar to the EGI champions programme, IDGF-SP are hoping to encourage the promotion of desktop grids usage inside and outside scientific organsitions. One of their current ambassadors is Professor Stephen Winter from the University of Westminster. Last year, the university saved £500,000 from deploying a desktop grid (DG).

IDGF-SP was launched in December 2012. Another aim is to build a desktop grid virtual team by:

• Promoting the technology in the EGI communities
• Adding DG resources to more virtual organisations (VOs)
• Collecting spare capacities from VOs
• Running new applications on the integrated infrastructure
• Finding EGI champion/IDGF ambassadors

The project will collect data in an application super-repository where users can see existing applications and associated metadata (i.e. attributes and implementations etc) and administrators have access to SZDG technical wiki (a consolidated knowledge base for DG related technologies).

Check out our educational website, Volunteer Garage (www.volunteer-garage.org). 

More than just computers: it's science! Meryam Tahar talks to e-ScienceTalk

Meryam Tahar is a year in industry student at the STFC. She tells us about her perspective on the Community Forum and why more young people should consider science and technology careers:


Thursday, April 11, 2013

Another EGI Champion interview: Fotis Psomopoulos

Fotis Psomopoulos develops data-mining algorithms for genomic research and protein modelling, and he is talking about life sciences:

Climate modelling on the grid: Eleni Katragkou, EGI Champion

Eleni Katragkou is a climate scientist from the University of Thessaloniki in Greece. She uses the grid to build more accurate climate models that integrate data at different scales, and is an EGI Champion. We managed to catch her during a break and this is what she had to say:


Using the grid to help make decisions in water network engineering

It's been unusually dry, for the first half of the week at least, for a conference in Manchester. There wasn't much water falling from the sky.

The community forum is here for people doing scientific computations, or computational science. Computer networks rely on infrastructure and network engineers are the experts in setting them up.

What about water distribution networks? (See what I did there? – a somewhat obtuse segueway, I grant you.) These marvels of 19th century engineering in the UK can present a challenge when it comes to managing them, mainly because you can't actually put sensors everywhere you might need them to figure out how flows can be regulated. Could you use grid computing and the information from sensors you can access to predict how the network is operating in places where sensing is impossible?

Well, yes!

John Brooke from Manchester University spoke to us about his work developing the architecture to enable this kind of prediction:

Silvia Olabarriaga discusses bioscience gateways

Silvia Olabarriaga accepted to talk about setting up a gateway for bioscience