Pages

Showing posts with label standards. Show all posts
Showing posts with label standards. Show all posts

Tuesday, October 30, 2012

One significant step for standards, a massive move forward for metadata

Metadata is data about data. At the EUDAT 1st conference in Barcelona, it was perhaps the most fêted buzz-word around, and for a very good reason. From a technological standpoint, the data tsunami can only be prepared for by being able to organize those vast quantities of data into manageable chunks, which means ‘tagging’ it so it can be referred to, searched for and easily accessed in the future.

When delicious.com (or, to give it its ‘proper name’ for stubborn stalwarts like me, del.icio.us)[1] arrived in 2003, the Web was already growing at an unbelievable rate. In 2000 there were a billion pages; by 2008, five years after del.icio.us’s inception, there were a trillion. Toolmakers, eager to solve problems in this age as in any other, were quick to provide solutions. Alongside improved search algorithms provided by Google and others, individual users of del.icio.us could curate their own ‘Web travelguide’, saving and signposting points and pages of interest by ‘tagging’ them however they liked – perhaps in categories relating to their hobbies, what they found funny, or were passionate about. The more socially-minded would carefully choose the tag words so that others could find them, and in this manner ‘social bookmarking’ was a major leap forward, building some of the foundations of the Web 2.0 era.
[1] For the record, I also believe box.net was better than box.com, despite it actually being the same thing.

In the same way, scientists, data scientists and e-infrastructure engineers have been thinking hard about how to add value to data by making it useful to others in the future. Data should be tagged to make it findable. But exactly how should it be tagged? Tags have to be flexible and dynamic to reflect the unpredictable nature of scientific research, but they have to follow standards, otherwise they’re little use to anyone else. How many of us at our first attempt at implementing a filing system end up with lots of similarly-named folders, each containing a single item, perhaps accompanied by a bulging folder called ‘misc’? Without standards developed with the experience of people who work with information and its management – librarians who have embraced the digital age – big data could end up being an incoherent, unwieldy mess, just like those first forays into filing. It’s perhaps not so ‘much catch-the-wave’ as ‘avoid the sea spray’ – all while a tsunami looms on the horizon.

“Without supporting tools, data isn’t data,” said Ross Wilkinson of the Australian National Data Service in the afternoon session on metadata; “—it rots! It needs to be made available [through e-infrastructures and computing resources] and it needs to be enhanced by making in available in alignment with other datasets.”
Metadata marsupial, the possum. (CC-BY Wolombi, Flickr)
Expanding on this point, Wilkinson explained that enhancing data through metadata allows curation of datasets with that real cross-disciplinary benefit, “not just to answer questions, but to explore the data to find new patterns”. By placing data in a rich (and that means metadata-loaded) context, scientists in Australia have been using habitat maps of where possums (which, we were told, don’t sweat) live to predict the likelihood of bush fires. Without those sweat glands the possums would not abide in areas likely to burn when the dry season comes. Finding this connection has profound economic and social implications for human habitations, construction and related policies in Australia, but before the data was curated and made open, that link might have never been found.

One area that definitely needs a robust approach to metadata is medicine, not only because of the rich terminology of biology, but because clinicians often like to see information in a diagrammatical format. This has presented problems for clinical metadata, because it’s harder to grep in a graphic than in a text file. Bernard de Bono of the European Bioinformatics Institute presented one solution, ApiNATOMY, which automates the creation of standard anatomy schematics and metabolic maps and allows the inclusion of metadata. It’s the standardization of the approach that those behind the project makes it suitable for multiscale anatomy analytics.

But what language will metadata be in? EUDAT is a project concerned with European data infrastructure, so it could be any one of 23 official languages recognised by the EU. Speaking in the plenary, Director of the Finnish IT Centre for Science, Kimmo Koski revealed that the standard language agreed on would be English – an important step towards European data standards.

Wednesday, March 16, 2011

SIENA Roadmap Working Session “How can eGovernment & DCI projects collaborate for mutual benefit”


Martin Walker and John Borras chaired the SIENA Roadmap Working Session “How can eGovernment & DCI projects collaborate for mutual benefit”. The panel consisted of Ignacio Blanquer, Tim Cowen, Evangelos Floros, Dawn Leaf, Steven Newhouse, Ian Osborne and Alan Sill. There are some statements I want to highlight here.

  • We need to identify how the existing e-Infrastructures can scale out to other sectors like Government.
  • Encourage people to experiment more. At the moment we use excuses for not doing anything, the most prominent is security.
  • Implementation and deployment need to take place in parallel with the development of standards. Of course there is a level of uncertainty before the development of standards but having a taxonomy and a reference architecture is a good starting point.
  • End-users are much more flexible than we think. We have to introduce more of end-user influence.
  • How do we integrate cloud in the existing standards like ITIL etc.?
  • One of the reason why clouds rely on standards is because clouds are new.

Tuesday, September 15, 2009

Steven Newhouse on EGI and middleware

I recently interviewed my fellow blogger Steven Newhouse, the interim EGI director, about how EGI has been coming along.

You can read his interview in iSGTW, but if you want to hear steven chat about the unversal middleware distribution within EGI have a listen to the clip below.


Monday, January 19, 2009

Cloudscape Conclusions: Topics in 2014

Cloudscape Mini Report - Back to the Future!
What will be the main topics of discussion in five years from now?

11 key predictions
  1. We will still be wrestling with regulatory issues, policy, compliance & security!
  2. When informing governments on the most important regulations, we should focus on regulatory standards.
  3. Carbon Issues!
  4. More control over the network! Regulatory issues are key to avoid "cloud wars".
  5. Complexities of large-scale distributed computing - they will be more network centric!
  6. Protecting information, unless we tackle compliance issues.
  7. Transferring simple jobs to complex existing applications & thinking of redesigning everything to move to the cloud, something IT departments are not keen on.
  8. Scale-out & Efficiency - somebody will have to redesign the application.
  9. Scaleability - Being able to be scaleable with multi-core machines or architectural machines.
  10. Industry-led standards: we set the standards that have been set by industry. Working in a reactive way.
  11. Continued discussions on the EC's Code of Conduct on Data Centres.

Friday, January 16, 2009

Some Farewell Thoughts

Back to sunny Pisa (far better running weather) with some final thoughts on the engagement of enterprises in the drive towards standards.

Siada El Ramly, Secretary General of the European Software Association, highlighted how cloud computing has the potential to bring major benefits to ISVs, especially in Europe. Key concerns are privacy and security, issues Ian Osborne (OGF-Europe and OGF's VP for Enterprise) says we will still be debating in comes to come.

What remains clear is the need for industry-led standards to address these issues and shine light on the benefits of doing so.

Thursday, January 15, 2009

Session 1 Cloud and distributed computing

A set of interesting talks on clouds, how they relate to earlier paradigms, how they will be exploited in the future. There was though the assertion that its too early for standards... but really sure if I agree since we have both open and closed source examples of the Amazon WS interfaces so are they becoming defacto???
It was also interesting that they mostly talked about Infrastructure as a service clouds (AWS, Eucalyptus) rather than all of the types, PaaS and SaaS examples glossed over.

Enric from the EC also mentioned how behind the work of OGF they are in terms of interoperability, which was very nice to hear. This means though that as far as we can see that NSF is the only large scale funding body that doesn't highlight this in its calls for funding oppurtunities.

Friday, June 6, 2008

Building Bridges between China & Europe

As part of my plan to explore international co-operation on grids, I had an enlightening chat with Gilbert Kalb, a senior scientist at Fraunhofer Gesellschaft in the Department for International Business Development. The interview helped shed new light on grid-enabled applications, interoperability and standards from an international perspective.

Gilbert currently manages Bridge, an EU funded project that connects EU and Chinese commercial & research organisations across several domains: simulation and design in aerospace industries; environmental disaster prediction; and drug discovery.

Gilbert, can you highlight a few key achievements of the Bridge project?
Bridge has achieved interoperability between major EU (GRIA) and Chinese (CNGrid GOS ) middleware. We have set up a platform for supporting applications and adopted a gateway with a high-level service based workflow approach to implement interoperability. Bridge has developed key technologies in both Grid middleware and in applications, including gateway-based interoperability, Master/Worker parallel programming model for grid computing, cross domain security policies mapping, and reliable massive data transfer.
The three applications all have different features, which in themselves bring a number of challenges: inter-continental workflow for optimisation; massive data transfer and processing for meteo-applications, and massive parallel computation for drug discovery. The key point is that they have all demonstrated the feasibility of Grid-enabled applications.

China has invested heavily in Grid research and in Grid infrastructure. It has not only a large market but can also offer great opportunities for research co-operation. What new experiences and knowledge have emerged from EU-China partnerships?
It is quite challenging to mange a project with partners from across Europe and China. We have to tackle not just technical problems, but also different time zones as partners operate tens of thousands of kilometres away. Technical issues tackled include providing partners with a seamless, efficient, reliable and affordable IT infrastructure, as well as full control of the usage of provided services and protection of intellectual property rights (IPR). In addition to these challenges, there are cultural differences and the need to communicate in a language that is often not our own.
But it is an interesting and valuable experience that has brought tangible results. We can now expect to see an increasing need to support large design and engineering projects on a global scale. My personal experience shows that all the partners and people involved in Bridge are prepared to go the extra mile: the recipe for success in an intercontinental research project. The strong relations established are paving the ground for further co-operation and have helped identify a number of potential partners from the aviation industry in China and the pharmaceutical industry.

What has been the value-add of attending OGF23?
OGF23 is an opportunity to demonstrate the findings of Bridge to a knowledgeable audience and discuss the various facets of the project with other people from the community. This kind of event also helps us understand the different stages of development and where we stand in the global community. Last but not least, this is a chance to network, make new contacts and pinpoint potential co-operations within Bridge.

How do you plan to engage with OGF over the coming months?
We are planning to make a contribution to standards in interoperability and grid-enabled application technologies by taking part in a number of international activities, including OGF working group activities.

Monday, June 2, 2008

OMII or not to OMII

The OMII-UK & OMII-EU session allowed these two flagship implementation and co-ordination projects from Europe show how far they had got in the creation of tools and software units that make use of OGF specifications. Shantenu Jha from LSU gave a really good presentation on the SAGA work that is being funding from OMII-UK. This tool is essential if we are going to move the use of grids and cyberinfrastrucutures out from the high level developers towards those area which build their own applications but are primarily domain researchers in e-Science or industry.

After the break the session continued with a description of the OGSA-DAI project from OMII-UK. Mike Jackson gave the talk. Sergio Andrezotti then talked about GLUE 2.0 and the GLUEMAN tool. The key thing that was shown from this though was that tools and software would need to be instrumented to publish compatible info.
Steve McGough described the status of GridSAM which is a tool for presentation of a JSDL compatible interface to the outside world. (Since I work on this project then I must declare an interest :)) The newest developments were described. Moving onto the original questions from Neil further examples of usage of the product were given.

All in all good to see people talking about actual implementations of standards as well as the support for getting useful standards through the OGF process!!