Friday, April 09, 2010

Genius for gene networks

Now that I cleaned the matrix data (see previous three posts) I can display them with Genius (google code) without having to deal with long genes descriptions:

First steps with Gridworks (3)

Once again I go back to the original file and, as I know that the first column has always the same format I will try: value.split(":")[-1] and the results will be:

And here we go:


In this case I've created a new column but I could have selected 'edit cell -> transform' to transform the original column directly.

Finally, to create exactly what I wanted:
value.split(':')[-2]+':'+value.split(':')[-1]

or
value.replace('Affymetrix:CompositeSequence:','')
or
value.split(':')[2,4].join(':')


In summary, with the last line of code, in a few seconds total (load the file and run the command) I can perform the desired transformation. Also, it would be helpful to be able to use the command "split multi-valued cells" with the choice of splitting into rows (as it is already possible now) or into columns (which is currently not possible).

First steps with Gridworks (2)

I am back on the 'Genius matrix' project and I notice that the column name is a bit cryptic, I therefore provide a new simple name.


Even if that was not the goal of my data cleaning I am still curious on how to split the first column in multiple ones using the separator ':' as criteria.


I select 'add column based on this column' and I get the following screen:

Using the "Gridworks Expression Language" (GEL) I can create a new column where I got rid of the prefix "Affymetrix:CompositeSequence" (as you can see in the above pic):

if(value.startsWith ("Affymetrix:CompositeSequence"), value.substring(29), value)

And the result is:


I still haven't used the separator ':' with my rule but the result is really close to what I wanted... and there is always the option to roll back.

Thursday, April 08, 2010

First steps with Gridworks (1)

I am having the chance of testing the alpha of Freebase Gridworks and given my appreciation for the work of David Huynh and Stefano Mazzocchi I am really excited about it. As these days I am working on a little software for visualizing gene networks and I need to perform a few simple steps for cleaning the data I decided to go Gridworks (video1 and video2). Gridworks is easy to install and once I run it, Gridworks opens a page in my browser.

I am then going to load a matrix in csv format simply filling in the data file location and a project name. Creating the project I get back the view of the matrix:


I now need to clean the first column that contains names such as:
Affymetrix:CompositeSequence:Rat230_2:1389668_at

and change them into something like: Rat230_2:1389668But first I want to visualize a larger number of rows (from 20 to 50) to make sure that the pattern is always the same. I select page size: 50 and this time the operation takes a few seconds to complete.
Now I want to find the best way to modify the first column and I start exploring the contextual menu:

Several quick options are available: transform to uppercase, lowercase, title case, collapse white spaces and so on. I also see a split multi-valued cells and this gets my attention. I select it, I specify as separator the semicolumn ':' and I get back:

This is not exactly what I wanted as I was expecting to get multiple columns out of the first one instead of getting multiple rows. Well, no problem the undo feature is one click away. You can either click on the top of the screen for the last command undo or you can browse the command history on the right of the screen and go back the desire amount of steps.


And I can go back to the original state.

Monday, February 22, 2010

Thursday, October 29, 2009

Heat-shock Proteins Families

Heat-shock proteins are named according to their molecular weight (kilodaltons):

Approximate molecular weight (kDa) Eukaryotic proteins Function
10 kDa Hsp10 Co-factor of Hsp60. Chaperonins: protein complexes that assist the folding of these nascent, non-native polypeptides into their native, functional state.
20-30 kDa The HspB group of Hsp. Ten members in mammals including Hsp27 or HspB1
40 kDa Hsp40 Co-factor of Hsp70
60 kDa Hsp60 Involved in protein folding after its post-translational import to the mitochondrion/chloroplast
70 kDa The HspA group of Hsp including Hsp71, Hsp70, Hsp72, Grp78 (BiP), Hsx70 found only in primates Protein folding and unfolding, provides thermotolerance to cell on exposure to heat stress. Also prevents protein folding during post-translational import into the mitochondria/chloroplast.
90 kDa The HspC group of Hsp including Hsp90, Grp94 Maintenance of steroid receptors and transcription factors
100 kDa Hsp104, Hsp110 Tolerance of extreme temperature

The Hsp70/Hsp40 Family (Chaperone)

The 70 kilodalton heat shock proteins (Hsp70s) are a family of highly-related protein isoforms ranging in size from 66 kDa to 78 kDa. Proteins with similar structure exist in virtually all living organisms. The Hsp70s are an important part of the cell's machinery for protein folding, and help to protect cells from stress. The various family members are distributed throughout different intracellular compartments but nevertheless share many structural and biochemical properties.

These include: the cytosolic/nuclear Hsp70 proteins Hsc70 (a.k.a. Hsp73) and Hsp70 (a.k.a. Hsp72), Grp78 or Bip present within the lumen of the endoplasmic reticulum, and Grp75 (also called mortalin) localized within mitochondria. Hsc70 is constitutively expressed and poorly stress-inducible, whereas Hsp70 is unabundant in normal physiological situations and strongly induced under oxidative stress. Hsc70 was placed in the heat shock protein family due to homology with other heat shock proteins.

All of the mammalian Hsp70 family members require one or more cochaperones for their reaction cycle (HSP40 or HSP100).

The Hsp60/Hsp10 Family (Chaperonins)

The HSP60 (GroEL - Prokaryotic) proteins in combination with its particular co-factor Hsp10 (GroES - Prokaryotic) bind newly synthesized polypeptides and facilitate their folding to the native state in an ATP-dependent cycle. The binding and sequestration of the substrate polypeptide occurs within the large central cavity of the chaperonin complex. It is thought that protection of the substrate protein within the central cavity of the chaperonin provides a sequestered protein folding environment, thereby reducing the probability of misfolding and aggregation of the target protein with other polypeptides. The chaperonins together with Hsp70 chaperones coordinate the efficient folding and assembly of many proteins throughout the cell.

Hsp60 families are highly immunogenic proteins. The related GroEL proteins from different pathogens elicit strong humoral and cellular immune responses.

In mammalian the Hsp60 is localized within mitochondria and with its cochaperone, Hsp10, participates in the folding and assembly of newly synthesized proteins as they are transported into the mitochondria from the cytosol.

Hsp90

Hsp90 is a molecular chaperone and is one of the most abundant proteins expressed in cells. Unlike some of the other well characterized heat shock proteins whose chaperone role involves their interaction with many cellular proteins, Hsp90 exhibits some selectivity for a distinct set of “client” proteins. Hsp90 interacts with a variety of protein kinases and transcription factors important for growth and development. Working with its very large number of cochaperones, Hsp90 appears to maintain its client proteins in a conformation that allows for their subsequent activation in response to appropriate growth signals. Not surprisingly, Hsp90 and its co-chaperones are at the forefront of research for those studying signal transduction events and cancer.

Sources

Heat Shock and Heat Shock Proteins

Heat Shock

In biochemistry, heat shock is the effect of subjecting a cell to a higher temperature than that of the ideal body temperature of the organism from which the cell line was derived. In fish that survive at 0°C, heat shock can be induced with temperatures as low as 5°C, whereas thermophilic bacteria that proliferate at 50°C will not express heat shock proteins until temperatures reach approximately 60°C. The process of heat shocking can be done in a CO2 incubator, O2 incubator, or a hot water bath.

Induction of heat shock is a method by which genes can be introduced into cells via a vector. This is done by mixing the vector with competent bacteria in a microcentrifuge tube. First, the tube is cooled to a low temperature for several minutes, usually with an ice bath. The tube is then quickly moved into warm water, preferably around 42°C. This sudden change in temperature causes the pores to open up to larger sizes, allowing DNA molecules to enter. After a brief interval, the tube is quickly cooled to a low temperature again. This closes up the pores, and traps the DNA inside. With this, the cells would have been transformed. However, as with almost all transformation techniques, this method is far from 100% efficient.

from Wikipedia (accessed on 10/28/2009)

Heat Shock Proteins

Heat shock proteins (HSPs) are present in cells under normal conditions (1), but are expressed at high levels when exposed to a sudden temperature jump or other stress (2).

(1) HSPs act like ‘chaperones,’ making sure that the cell’s proteins are in the right shape and in the right place at the right time. For example, HSPs help new or distorted proteins fold into shape, which is essential for their function. Heat shock proteins also shuttle proteins from one compartment to another inside the cell, and transport old proteins to ‘garbage disposals’ inside the cell. Heat shock proteins are also believed to play a role in the presentation of pieces of proteins (or peptides) on the cell surface to help the immune system recognize diseased cells. HSPs appear to serve a significant cardiovascular role.

Chaperones are proteins that assist the non-covalent folding/unfolding and the assembly/disassembly of other macromolecular structures, but do not occur in these structures when the latter are performing their normal biological functions. Chaperones do not necessarily convey steric information required for proteins to fold: thus statements of the form `chaperones fold proteins` can be misleading. One major function of chaperones is to prevent both newly synthesised polypeptide chains and assembled subunits from aggregating into nonfunctional structures. It is for this reason that many chaperones, but by no means all, are also heat shock proteins because the tendency to aggregate increases as proteins are denatured by stress. However, 'steric chaperones' directly assist in the folding of specific proteins by providing essential steric information.

Heat shock proteins trigger immune response through activities that occur both inside the cell (intracellular) and outside the cell (extracellular).

  • Intracellular activities: Because of the normal functions of heat shock proteins inside the cell (such as helping proteins fold, preparing proteins for disposal, etc.), HSPs end up binding virtually every protein made within the cell. This means that at any given time, HSPs can be found inside the cell bound to a wide array of peptides that represent a ‘library’ of all the proteins inside the cell. This library contains normal peptides that are found in all cells as well as abnormal peptides that are only found in sick cells. Research suggests that inside the cell, heat shock proteins take the peptides and hand them over to another group of molecules. These other molecules take the abnormal peptides that are found only in sick cells and move them from inside the cell to outside on the cell’s surface. When the abnormal peptides are displayed in this way, they act as red flags, warning the immune system that the cell has become sick. These abnormal peptides are called antigens — a term that describes any substance capable of triggering an immune response.
  • Extracellular activities: Heat shock proteins are normally found inside cells. When they are found outside the cell, it indicates that a cell has become so sick that it has died and spilled out all of its contents. This kind of messy, unplanned death is called necrosis, and only occurs when something is very wrong with the cell. Extracellular HSPs are one of the most powerful ways of sending a ‘danger signal’ to the immune system in order to generate a response that can help to get rid of an infection or disease.

(2) HSPs are a part of the cell's internal repair mechanism. They are also called stress-proteins. They respond to heat, cold and oxygen deprivation by activating several cascade pathways. In fact, HSPs stabilize proteins and are involved in the folding of denatured proteins. High temperatures and other stresses, such as altered pH and oxygen deprivation, make it more difficult for proteins to form their proper structures and cause some already structured proteins to unfold. Left uncorrected, mis-folded proteins form aggregates that may eventually kill the cell. HSPs are induced rapidly at high levels to deal with this problem. Increased expression of HSPs is mediated at multiple levels: mRNA synthesis, mRNA stability, and translation efficiency.

Heat shock factor 1 (HSF-1) is the major regulator of heat shock protein transcription in eukaryotes. In the absence of cellular stress, HSF-1 is inhibited by association with heat shock proteins and is therefore not active. Cellular stresses, such as increased temperature, can cause proteins in the cell to misfold. Heat shock proteins bind to the misfolded proteins and dissociate from HSF-1. This allows HSF1 to form trimers and translocate to the cell nucleus and activate transcription.

from Wikipedia (accessed on 10/28/2009), from Wikipedia (accessed on 10/28/2009), from Wikipedia (accessed on 10/28/2009), from Wikipedia (accessed on 10/28/2009)


Friday, October 23, 2009

Eclipse+Grails+GWT: Integrating 'live' GWT [1]

After creating a Grails project in Eclipse (and a plugin), I want now to make use of GWT (Google Web Toolkit). For doing so I am going first to install the useful Google plugin for Eclipse (download). As I am working with Eclipse Galileo, I am installing the software using: http://dl.google.com/eclipse/plugin/3.5

If you are new to GWT I would suggest you to start creating a simple GWT project (without any Grails involved). Follow this.

Adding the gwt libraries to the classpath of the Grails project
If you installed the Eclipse plugin for GWT this would be a fairly simple step. Right click on the project and Properties->Java Build Path->Libraries->Add Library and I will find in the list the Google Web Toolkit in the list (this is obviously because I previously succesfully installed the Eclipse plugin for GWT).

An alternative way (to the Google plugin) would have been to define in the Eclipse Preferences->Java->Build Path->Classpath Variables the variable GWT_HOME that will point to the gwt installation (i.e. /Users/paolociccarese/Library/gwt-mac-1.7.1). If you haven't installed GWT yet, you can download it here. Then we are going to add to the libraries of the 'FirstApp' projects the jars included in the GWT distribution (Right click on the project and Properties->Java Build Path->Libraries->Add variable you can select GWT_HOME and extend it with the jars: gwt-api-checker, gwt-dev-mac, gwt-servlet, gwt-user.

Another way would be using ivy for resolving the dependencies.

The Grails plugin for GWT
Then I will install the Grails plugin for GWT. Once again, after positioning in the root of the 'FirstApp' project I will type:
> grails install-plugin gwt

At this point I am almost ready to go...



Monday, October 19, 2009

Eclipse+Grails: Adding a 'live' plugin to the project

Once created the new project 'FirstApp' I want to be able to create a Grails plugin in the Eclipse IDE. To create a plugin is pretty straightforward, I use the command line, after positioning in the root of the current workspace, and I type:
> grails create-plugin FirstPlugin

The Grails plugin structure is almost identical to the Grails application structure. See my previous post about the Grails plugin. I am going to import in Eclipse the plugin with the classic 'Import as Existing Project'. As I did for the 'FirstApp' I can enable the Groovy capabilities right clicking on the project and selecting: Configure-> Convert to Groovy project. Also, exactly as I did for the Grails application, I select the ivy.xml file, I right-clik and select 'Add Ivy Library...'. This will link, through ivy, all the libraries necessary to the newly created Grails application to compile and run. Once performed this step I can safely remove all the libraries linked in the subdirectory GRAILS_HOME from the Properties->Java Build Path. You should also notice a new icon in the toolbar with tooltip 'Resolve All Dependencies'. You can use that icon to retrieve new dependencies that you specified in the ivy.xml file.

For testing purposes I can run the plugin as it was an application with:
> grails run-app

After checking that the plugin can run I want to use the Grails plugin 'FirstPlugin' in the Grails application 'FirstApp'. The 'classic' way would be to package the plugin by command line. After positioning in the root of the plugin:
> grails package-plugin
will create a file like grails-firstplugin-0.1.zip. After positioning in the root of the app, I can then the installation of the plugin:
>grails install-plugin ../FirstPlugin/grails-firstplugin-0.1.zip

Unfortunately, this approach is really annoying and time consuming when using the Eclipse framework. In fact, every time the plugin is changing I have to package it again and install it again.

Well, in Eclipse there is a better way to go. First of all we create the file conf/BuildConfig.groovy in the 'FirstApp' project. In this file we are going to type:
grails.plugin.location.firstPlugin = "../FirstPlugin"

This will basically tell Grails to refer to that plugin specifying the plugin folder. To test this, after creating the file, you can proceed creating a controller in the plugin. After positioning in the plugin root, I type:
>grails create-controller org.example.First

This will generate a new controller
org.example.FirstController .To prove that the BuildConfig.goovy file is working you can run the application. After positioning in the root of the application:
>grails run-app

Checking the URL: http://localhost:8080/FirstApp/ you should see the controller (of the plugin) showing up. There is still a problem though if you create a second controller (or a change in the code of the plugin), this will not show up in the application unless you restart the server. This might become extremely annoying on the long run.

The solution to that is the creation of a soft link. We are going to create a link to the plugin project inside the code of the application. First I create the directory plugins in the FirstApp project:
>mkdir plugins

Then, in that directory, I am going to create the link to the FirstPlugin:
>ln -s ../../FirstPlugin

Right. Now, if you followed correctly (...and if I have not forgot any step), while the server is running, you should be able to create a new controller in the 'FirstPlugin' and see it in the list of the controllers in the running application: http://localhost:8080/FirstApp/

One more thing, remember to refresh the eclipse project every time you make use of a Grails command that creates/modify files.

Eclipse+Grails: Setting up a new project

Here is my current setup for creating a Grails project with Eclipse:
  1. Installed Eclipse jee 3.5 Galileo SR1 (download)
  2. Installed SVN plugin Subclipse (Update site: http://subclipse.tigris.org/update_1.6.x)
  3. Installed Groovy plugin (Update site: http://dist.codehaus.org/groovy/distributions/greclipse/snapshot/e3.5/)
  4. Installed Ivy plugin (http://www.apache.org/dist/ant/ivyde/updatesite)
I am then going to setup the launcher for the Grails commands through the "External Tools Configuration". I am specifying the 'location of Grails" (i.e. /Users/paolociccarese/Library/Grails-1.1.1/bin/grails), then ${project_loc} as 'Working directory' (that means the command will be related to the selected project in the Eclipse workspace. Finally I define as 'Arguments': ${string_prompt} . This will allow me to run the Grails commands like I would do from the command line. To be honest I also use the command line when I am managing grails projects, for some operations is more practical.

Finally I create the GRAILS_HOME variable in Eclipse->Preferences->Java->Build Path->Classpath Variables. This is used by Eclipse to compile the project. You may notice that GROOVY_HOME and IVY_HOME have already been created by the plugins.

Now I want to create a new Grails application that I will call 'FirstApp'. For doing this I will use the command line and, after positioning myself in the root of the current workspace, I type:
> grails create-app FirstApp

This will create the structure of my new application (in the newly created directory: FirstApp). I can now import this structure as 'Import as Existing Project' in my workspace. I can enable the Groovy capabilities right clicking on the project and selecting: Configure-> Convert to Groovy project.

Now if you right-click on the project and select: Properties->Java Build Path you will notice tha all the libraries are linked in a subdirectory of GRAILS_HOME. This is not my solution as I want to be dependent from the ivy.xml file. To do so, I select the ivy.xml file, I right-clik and select 'Add Ivy Library...'. This will link, through ivy, all the libraries necessary to the newly created Grails application to compile and run. Once performed this step I can safely remove all the libraries linked in the subdirectory GRAILS_HOME from the Properties->Java Build Path. You should also notice a new icon in the toolbar with tooltip 'Resolve All Dependencies'. You can use that icon to retrieve new dependencies that you specified in the ivy.xml file.

Now I can test if the application is actually running. I select a file in the project and I run the external application 'Grails' that I previously created. A window will pop up and we can fill out the text box with the command 'run-app'. The alternative would have been to use the shall and, after positioning in the root of the project, type:
> grails run-app

If everything went right we can now browse: http://localhost:8080/FirstApp/ getting back the Grails welcome screen.

Wednesday, September 16, 2009

Uploading RDF vocabularies (GRAILS, Virtuoso and Sesame)

If you need to upload an RDF vocabulary in Virtuoso, the easiest way to go is doing it through the Virtuoso Conductor. It is pretty straightforward, you define the graph name and the file to upload and that is it. When developing an application though, you may want to be able to upload the files programmatically (for instance for managing metadata associated to the vocabulary and creating a catalog).

The following snipped of code is a modified version of the one I found in the openrdf forum:

Repository myRepository = new VirtuosoRepository("jdbc:virtuoso://localhost:1111","dba","dba");
myRepository.initialize();
File file = new File("/Users/paolociccarese/Desktop/contact.rdf");
RepositoryConnection con = myRepository.getConnection();
URI context = new URIImpl("http://paolociccarese.info/contact");
con.add(file, null, RDFFormat.RDFXML, context);
con.close();
For running this code it is necessary to perform a bunch of imports:

import org.openrdf.model.URI;
import org.openrdf.model.impl.URIImpl;
import org.openrdf.rio.RDFFormat
import org.openrdf.repository.Repository
import org.openrdf.repository.RepositoryConnection
To make it work, if you use grails and ivy (ivy.xml), you should add at least the following dependencies:

<dependency org="org.openrdf" name="openrdf-repository-api" rev="2.0.1" conf="compile"/>
<dependency org="org.openrdf" name="openrdf-rio-rdfxml" rev="2.0.1" conf="compile"/>
<dependency org="org.openrdf" name="openrdf-model" rev="2.0.1" conf="compile"/>

and you also need to link to the maven repository for these in the ivysettings.xml file:

<ibiblio name="aduna" root="http://repository.aduna-software.org/maven2" m2compatible="true"/>

and finally run the command >grails get-dependencies again.

Friday, September 11, 2009

Grails: Multiple datasources without GORM

Who's playing around with Grails knows about GORM. GORM is Grails' object relational mapping (ORM) implementation. GORM is built on the top of Hibernate and enables persistence. The configuration of the database used for persistence is in the file grails-app/conf/DataSource.groovy. The application might need to access other databases for reasons other than persistence. It might be necessary to distinguish different sources for development, test or production (the Grails standard environments). In this case it is possible to leverage Spring creating the file grails-app/conf/srping/resources.groovy with a content similar to the following, which is setting up the connection to the Virtuoso triple store:

import org.apache.commons.dbcp.BasicDataSource

/*
* @author Paolo Ciccarese
*
* The virtuoso data source has been defined here to not
* interfere with the GORM datasource in the file
* grails-app/conf/DataSource.groovy
*/

beans = {
// Virtuoso triple store
virtuosoDataSourceDevelopment(BasicDataSource) {
driverClassName = "virtuoso.jdbc3.Driver"
url="jdbc:virtuoso://localhost:1111"
username="dba"
password="dba"
}
virtuosoDataSourceTest(BasicDataSource) {
driverClassName = "virtuoso.jdbc3.Driver"
url="jdbc:virtuoso://blah.org:1111"
username="dba"
password="dba"
}
virtuosoDataSourceProduction(BasicDataSource) {
// Here the details of the production data source
}
}
BasicDataSource has been used as it guarantees connection pooling. In some examples, instead of that class you can find org.springframework.jdbc.datasource.DriverManagerDataSource . Reading the documentation: This class is not an actual connection pool; it does not actually pool Connections. It just serves as simple replacement for a full-blown connection pool, implementing the same standard interface, but creating new Connections on every call. This might create problems with Virtuoso in the case of numerous queries. Once we defined the data sources it is possible to define a service (> grails create-service) for providing the right one according to the current environment:

class SparqlService {

def transactional = false

def virtuosoDataSourceProduction
def virtuosoDataSourceDevelopment
def virtuosoDataSourceTest

private static final String TEST = "test"

def getDataSource() {
if(grails.util.GrailsUtil.isDevelopmentEnv())
return virtuosoDataSourceDevelopment
else if(grails.util.GrailsUtil.getEnvironment()==TEST)
return virtuosoDataSourceTest
else
return virtuosoDataSourceProduction
}
}
Now it will be possible, for other services to use the one just defined. Something like:

class SparqlQueryService {

def sparqlService

boolean transactional = false

public List<Map<String, Object>> getSparqlResultRowMaps(String query){
log.info("Running query: " + query);
JdbcTemplate jdbcTemplate = new JdbcTemplate(sparqlService.getDataSource())
.....
return results;
}

....
}

Where the sparqlService is injected. Thus when a query is needed is sufficient to refer to sparqlQueryService, which will take care of selecting the right datasource.

Friday, September 04, 2009

The Grails plugin class

GrailsImage via Wikipedia

After creating the plugin structure one thing you can do is to edit the Grails plugin class. The grails plugin class, by default, looks something like this:

class UtilsGrailsPlugin {
// the plugin version
def version = "0.1"
// the version or versions of Grails the plugin is designed for
def grailsVersion = "1.1.1 > *"
// the other plugins this plugin depends on
def dependsOn = [:]
// resources that are excluded from plugin packaging
def pluginExcludes = ["grails-app/views/error.gsp"]

// TODO Fill in these fields
def author = "Your name"
def authorEmail = ""
def title = "Plugin summary/headline"
def description = '''\\
Brief description of the plugin.
'''

// URL to the plugin's documentation
def documentation = "http://grails.org/Utils+Plugin"

....
}
As we are writing our first plugin there are not dependencies we want to define. We will simply fill out some details:

def author = "Paolo Ciccarese"
def authorEmail = "paolo.ciccarese@gmail.com"
def title = "Utilities for vocabularies/ontologies management"
def description = '''\\
Utilities for vocabularies/ontologies management
'''
You might want also to define the location of the documentation. This file will be used to generate the plugin package when needed. It is sufficient to be in the root directory of the plugin and type:
> grails package-plugin

This will create a grails-utils-0.1.zip file that can be then installed by other grails plugins/applications. To install it will be sufficient to be in the new application/plugin root directory and to type
> grails install-plugin grails-utils-0.1.zip

But before doing this, you probably have to write some code that will provide actual utilities.



Reblog this post [with Zemanta]

A few mins with Grails (Creating a plugin)

GrailsImage via Wikipedia

Lately I found myself playing around with Grails and I've been appreciating right away the plugins mechanism (Grails website: creating a Plugin).

With a simple:
> grails create-plugin Utils

or even with a simpler (the script will ask you to specify the name of the plugin later on):
> grails create-plugin

you will be able to generate the structure of a Grails plugin that you will notice is really similar to the structure of a Grails application (in italic the directories):

.classpath -> Eclipse .classpath file. It includes all the references to the folders in the project that are going to be in the classpath and all the references to the libraries for making your grails project run. The libraries are expressed through the GRAILS_HOME variable under the form GRAILS_HOME/dist/grails-web-1.1.1.jar. The directories that are included in the classpath are typically (you can detect them easily in Eclipse if you import the project just created as "existing project"):
  • src/java
  • src/groovy
  • grails-app/conf
  • grails-app/controllers
  • grails-app/domain
  • grails-app/services
  • grails-app/taglib
  • test/integration
  • test/unit
.project -> Eclipse .project file
.settings -> Eclipse .settings directory

Utils-test.launch -> can be used to run the unit tests
Utils.launch -> can be used to run the application
Utils.tmproj

UtilsGrailsPlugin.groovy -> The plugin class defines the version of the plugin and optionally various hooks into plugin extension points.

application.properties -> Contains a few lines to describe the details of the plugin/application:
#utf-8
#Fri Sep 04 10:12:01 EDT 2009
app.version=0.1
app.servlet.version=2.4
app.grails.version=1.1.1
app.name=Utils
build.xml -> It is an Ant file that contains a set of targets, allowing to resolve dependencies declared in the Ivy file, to compile and run the sample code, produce a report of dependency resolution, and clean the cache or the project. Ivy is a subproject of Ant.
grails-app -> directory containing the main components of a Grails application/plugin: controllers, services... (see classpath)
ivy.xml -> Ivy file containing the description of the dependencies of a module. Ivy (Ivy website) supports dependency resolution (grails does not support Maven).
ivysettings.xml -> file for Ivy configuration
lib -> directory containing the list of libraries for running the application
scripts -> Grails scripts
src -> directory with the source codes in both java and groovy (see classpath)
test -> directory collecting all Grails unit tests: Grails supports the concepts of unit and integration testing. Unit testing is for small focused, fast loading tests that don't load supporting components. Integration testing is for tests that load the surrounding environment (it supports injection).
web-app -> contains the web components of the Grails project

Reblog this post [with Zemanta]

Wednesday, December 31, 2008

Moving towards the SWAN Collections Ontology [2]

The idea of reasoning with sequential structures in OWL-DL is appealing. However, as already mentioned, we cannot use the RDF vocabulary in OWL-DL.

Drummond et al. [1] proposed a way of representing sequential structures in OWL-DL. Analyzing the work of Hirsh & Kudenko [2] Drummond argued that "their representation requires extensive rewriting, the relation of the resulting structures to the original lists is not intuitive and, more importantly, the resulting structures grow as the square of the length of the list". Then, he describes a general list pattern, an intuitive approach related to that suggested by Hayes [3] and incorporated in the Semantic Web Best Practice Working Group’s note on n-ary relations.

The list pattern works as follow:
Each item is held in a “cell” (OWLList); each cell has 2 pointers, one to a head (hasContents - functional) and one to the tail cells (hasNext - functional); the end of the list is indicated by a terminator (EmptyList) which also serves to represent the empty list. A transitive property, isFollowedBy, as a super-property of hasNext as been defined as well. In other words the members of any list are the contents of the first element plus the contents of all of the following elements. A separate OWL vocabulary has been defined as the RDF vocabulary cannot be used in OWL-DL.

Through the transitive property followedBy it is possible to ask things like: give me all the items that are followedBy "AC" for instance and it doesn't matter what is in between the item and the sequence itself.

In Manchester Syntax:


OWLList can express:


For instance for the pattern (A*):

List_only_As --> List AND
hasContents ONLY A AND
isFollowedBy ONLY (List AND hasContents ONLY A)

Still, in OWL-DL there are a bunch of constraints that cannot be defined (and I would suggest to read the paper for the complete list).

The list ontology page.


[1] Nicholas Drummond, Alan Rector, Robert Stevens, Georgina Moulton, Matthew Horridge, Hai Wang and Julian Sedenberg (2006). Putting OWL in Order: Paterns for sequences in OWL. OWL Experiences and Directions (OWLED 2006), Athens, Georgia, USA.
[2] Hirsh, H. and D. Kudenko. Representing Sequences in Description Logics. in Fourteenth National Conference on Artificial Intelligence. 1997.
[3] Noy, N.F. and A. Rector, N-ary relations. 2004, Editors Draft, Semantic Web Best Practices Working Group, W3C.

Tuesday, December 30, 2008

Moving towards the SWAN Collections Ontology [1]

RDF Containers
RDF allows the usage of three kinds of containers:
  • rdf:Bag - A Bag represents a group of resources or literals, possibly including duplicate members, where there is no significance in the order of the members. For example, a Bag might be used to describe a group of part numbers in which the order of entry or processing of the part numbers does not matter.
  • rdf:Seq - A Sequence or Seq represents a group of resources or literals, possibly including duplicate members, where the order of the members is significant. For example, a Sequence might be used to describe a group that must be maintained in alphabetical order.
  • rdf:Alt - An Alternative or Alt represents a group of resources or literals that are alternatives (typically for a single value of a property). For example, an Alt might be used to describe alternative language translations for the title of a book, or to describe a list of alternative Internet sites at which a resource might be found. An application using a property whose value is an Alt container should be aware that it can choose any one of the members of the group as appropriate.
For example, a statement about "The resolution was approved by the Rules Committee, having members Fred, Wilma, and Dino" will have the form in triples:

ex:resolution exterms:approvedBy ex:rulesCommittee .
ex:rulesCommittee rdf:type rdf:Bag .
ex:rulesCommittee rdf:_1 ex:Fred .
ex:rulesCommittee rdf:_2 ex:Wilma .
ex:rulesCommittee rdf:_3 ex:Dino .

and this is much better than:

ex:resolution exterms:approvedBy ex:Fred .
ex:resolution exterms:approvedBy ex:Wilma .
ex:resolution exterms:approvedBy ex:Dino .

since these statements say that each member individually approved the resolution.

For further examples RDF Primer.

But containers only say that certain identified resources are members; they do not say that other members do not exist. There is no way to exclude that there might be another graph somewhere that describes additional members.

RDF Collections
RDF provides support for describing groups containing only the specified members, in the form of RDF collections. An RDF collection is a group of things represented as a list structure in the RDF graph.

in RDF/XML a collection is something like this:
<rdf:Description rdf:about="http://e.org/family/349">      
<s:familyMembers rdf:parseType="Collection">
<rdf:Description rdf:about="http://e.org/person/Paolo"/>
<rdf:Description rdf:about="http://e.org/person/Emanuele"/>
<rdf:Description rdf:about="http://e.org/person/Maria"/>
<rdf:Description rdf:about="http://e.org/person/Franco"/>
</s:familyMembers>
</rdf:Description>
this can also be written in RDF/XML by writing out the same triples (without using rdf:parseType="Collection") using the collection vocabulary:
<rdf:Description rdf:about="http://e.org/family/349">
<s:familyMembers rdf:nodeID="sch1"/>
</rdf:Description>

<rdf:Description rdf:nodeID="sch1">
<rdf:first rdf:resource="http://e.org/person/Paolo"/>
<rdf:rest rdf:nodeID="sch2"/>
</rdf:Description>

<rdf:Description rdf:nodeID="sch2">
<rdf:first rdf:resource="http://e.org/person/Emanuele"/>
<rdf:rest rdf:nodeID="sch3"/>
</rdf:Description>

<rdf:Description rdf:nodeID="sch3">
<rdf:first rdf:resource="http://e.org/person/Maria"/>
<rdf:rest rdf:nodeID="sch4"/>
</rdf:Description>

<rdf:Description rdf:nodeID="sch4">
<rdf:first rdf:resource="http://e.org/person/Franco"/>
<rdf:rest rdf:resource=
"http://www.w3.org/1999/02/22-rdf-syntax-ns#nil"/>
</rdf:Description>
For more examples RDF Primer.

RDF imposes no "well-formedness" conditions on the use of the collection vocabulary (it is possible, for instance, to define multiple rdf:first elements), thus, RDF applications that require collections to be well-formed should be written to check that the collection vocabulary is being used appropriately, in order to be fully robust. Maybe OWL which can define additional constraints on the structure of RDF graphs, can rule out some of these cases?

OWL and Ordering
OWL have no support for ordering, but the natural constructs from the underlying RDF vocabulary (rdf:List and rdf:nil) are unavailable in OWL-DL because they are used in its RDF serialization. In principle, rdf:Seq is not illegal but it depends on lexical ordering and has no logical semantics accessible to a DL classifier. In other terms: (1) The elements in a container are defined using the relations rdf:_1, rdf:_2, and so on that have no formal definition in RDF. Using them for the purpose of reasoning will require us to define and enforce the properties of these relations. (2) It is not possible to define a container that has elements only of a specific type. (3) For updating a specific element in a container in a remote source, one is forced to transmit the whole container. (4) It is not possible to associate provenance information with the elements in a container [1].

But OWL has greater expressivity than RDF (with constructs such as transitive properties) and reasoning capabilities (for checking consistency and inferring subsumption). Thus, the idea of reasoning with sequential structures in OWL-DL looks appealing.

[1] Vinay K. Chaudhri, Bill Jarrold, John Pacheco. Exporting Knowledge Bases into OWL. OWL Experiences and Directions (OWLED 2006), Athens, Georgia, USA.

Friday, November 21, 2008

SWAN Ontology v. 1.2 almost ready to go

In the last months, I've been busy in developing the new version of the ontology [SWAN Ontology] that represents the "backbone" of the SWAN project. In this iteration, I had two major goals in mind: modularity and provenance. The new ontology is composed by a set of modules, actually the SWAN ontology consists of a collection of ontologies. I think this is an important step for several reasons:
  • first of all the SWAN ontology is growing in size. Modules can help in managing the increasing complexity
  • modules can improve the learning process of people that want to approach our ontology for modeling scientific discourse or simply for reusing a part of it
  • defining modules helped me in thinking a little bit more

The architecture of the SWAN ontology release candidate

As additional feature I was also thinking to provide sub-modules for increasing reuse without asking potential users to write their own subset of the SWAN ontology. Thus, for instance, the Agents ontology is split in different modules that can include or not provenance and/or collections. This is because I assume that not everybody wants to deal with the tedious ordered lists and not anybody needs to define provenance the level we need.

Provenance is one of the major aspects in the semantic web world that we are trying to build with SWAN. Our application is mashing up data coming from different sources and we would like to be able to export the new knowledge product giving credit to the original data provider and declaring which piece of software performed the conversion of such data into our format.

I will speak more in detail about provenance in my next post.

[SWAN Ontology] Ciccarese P, Wu E, Kinoshita J, Wong G, Ocana M, Ruttenberg A, Clark T. The SWAN Biomedical Discourse Ontology. Journal of Biomedical Informatics, in press. PMID: 18583197

Wednesday, March 05, 2008

SWAN - Semantic Web Applications in Neuromedicine [2]

Thus, SWAN is not a like Wikipedia because several "hypotheses" (consistent or inconsistent) can co-exist. to be more precise I would say that the SWAN ontology is an ontology for modeling scientific discourse. Thus, I would define discourse elements as key entities in the SWAN ecosystem. They represent the hubs of the scientific discourse, or in general of the discourse.

Figure 1 - Walsh Hypothesis in the SWAN browser

Looking at fig. 1 it is possible to see the title of the Hypothesis, a description, the authors of such hypothesis (in this case the authors are the authors of the journal article the hypothesis has been derived from). Then, after the journal article used as source of the informatin related to the hypothesis we have the contained discourse elements. Right, a hypothesis can contain a list of discourse elements. In this case we have a list of claims (scientifically proved discourse elements) but it is possible to have in the discourse elements list other hypothesis, research questions or comments...

The SWAN Team: Tim Clark, June Kinoshita, Paolo Ciccarese, Marco Ocana, Gwen Wong, Elizabeth Wu.

Monday, March 03, 2008

SWAN - Semantic Web Applications in Neuromedicine [1]

In the last months, Marco and I have been coding for the SWAN project for Mass General Hospital (Neurology Dept) and Harvard Medical School. The SWAN project is the reason I moved to Boston to work. It is not easy to explain in a few words what SWAN (that stands for Semantic Web Applications in Neuromedicine) does (or it is supposed to do). I could say that 'we are using Semantic Web technologies with the idea of helping the researchers' life' but I understand that this is not really useful.

I'll try to explain it better with an example. When I was a student, I used to create summaries of the lessons integrating my notes with what I was finding in some books. I was using obviously (I am not that young anymore) paper sheets, colors, drawings... and so on. I had my formalism for stressing a definition, a theorem a short summary and whatsoever. It was efficient, I could easily remember the things (I have visual memory) and it was faster than going through the book again and again. This was perfect for a single lesson. But what was happening with an entire year of lectures? With different topics somehow connected each others? Well, I tried to update the things but it was hard and everything was getting terribly messy. At that time, the word processors were really poor and crispy. Now we can think of organizing the things in some electronic documents... better we can use a wiki where several students can cooperate to build faster with less effort. Everybody knows wikipedia right? Nice, we have wiki tools, we can decide our formalism, the meaning of the colors... this works if we have "one truth". Let's say I want to create in Wikipedia a page about a politician and I really dislike him/her (something that occurs often to me). I would probably be aggressive and biased. Somebody else could have a different perspective on the same person... this needs a mediation and rules to follow.

Now, the same perspective can be true in science. When we have hypothesis these are still not confirmed facts. Scientists need to prove them and it is normal to have disagreement. Disagreement is part of the scientific process (and as we are not in the Middle Ages we don't risk our life saying something 'different'... I guess). In science disagreement can be a real value.

SWAN is not Wikipedia, it is in some perspective the opposite of it. In SWAN, several 'truths' or better 'hypotheses' (consistent or not) can exist at the same time... inconsistencies can be both declared and inferred (nice uh?). In SWAN we can build the map of science (well, a part of it)... (TO BE CONTINUED)

Thursday, February 14, 2008

Classes which are things and classes 'about' things

One of the most interesting distinctions that I keep always in mind when I create an ontology is what is representing a 'real thing' and what is 'talking about a real thing'. Let's consider an example related to bioinformatics. I want to create an ontology which is modeling proteins. Nowadays there are different sources where we can find information about proteins. If we are building a system performing data integration, we probably don't want to copy all the data belonging to those sources in our knowledge base. It is more correct to build references, sort of records that are pointing to the original source when the user/system wants to know more. What we are building are records, entities that 'talk about' real things like proteins. Vice versa, if we want to provide content about proteins (i.e. providing proteins variants) we would probably model the real things, the actual proteins. It doesn't really make sense to say that a record 'hasVariant' another record. Maybe a record 'refersToVariantRecord' or something like that.

But why all this? Well, this is helping in building the models. Let's say that I want to define a 'authoredBy' property for a scientific article. Now, if I consider the real thing (i.e. the actual article) I can say 'authoredBy' but if i am building a record of the article (a reference like the ones that PubMed does) and I say 'authoredBy' am I referring to the article or to the record? As ontologies are meant to define semantic.... I guess this is a crucial point.