29 December 2016

Omeka vs Heurist: an historian's perspective

Archaeologists and historians share a concern over the dimension of space and the time-space binary. The study of landscapes, a construct that is defined by this binary, has thus become central to both disciplines. My own interest in landscapes and space in the study of the past is what has brought me to this archaeological conference.

I am presenting two on-going digital projects, one on landscapes and the other on a census from the early years of Greek statehood in the nineteenth century; both with a marked spatial component. I will try to explain my selection of web-based platforms, which took into account the design features of each platform and their suitability for the particular aims of each project. The two platforms in question are Omeka of the  Roy Rosenzweig Center for History and New Media and Heurist, the brain-child of Ian Johnson of the Faculty of Arts and Social Sciences of the University of Sydney. Both are free, opensource and promise sustainability and are therefore suitable for the underfunded pursuits of historians. Heurist moreover offers free hosting at the Sydney University data center. They require little expert IT help in contrast to custom-made solutions that often sustain a habit of reinventing academic wheels. Although they are fundamentally different in design, they do offer some similar functionalities of interest to both projects, such the ease to create maps and timelines and to insert semantic links.

The choice was defined by the structure inherent to the set of data and, no less important, by the intention of each project. As far as the landscapes project is concerned, my initial motivation was to publish a number of historical landscape photographs which were so poor in quality that no publisher would touch, the assumption being that the internet is much more forgiving. In its present shape the project consists of a compilation of textual and pictorial representations of Greek landscapes. Landscapes are perceived as “a human–environmental interactive sphere, transforming over time”, which are shaped both conceptually and ecologically by the cultural interaction among humans and by evolutionary transformations that also involve other species, and constitute places “upon which past events have been inscribed ..., on the land. The purpose of the project as it developed eventually is to generate narrations based on the  historical texts, for instance travellers’ accounts, images and, hopefully in the future, oral accounts. These narrations may be mere reconstructions of past travellers’ experiences but may also involve the evolution of particular sites as landscapes over time, much in the form of a palimpsest, or may interpret changes observed on specific types of environment under ecological and/or social pressures, for instance aquatic landscapes, and thus formulate and elaborate on historical questions. For practical purposes, I begin with a manageable number of  landscapes studied in my book on the history of malaria.

In other words, in constructing narratives, the project is selective, experiential  and is based on the interpretation of evidence, whose compilation is linear in structure. I will explain why I preferred Omeka as a tool.

Omeka uses the Extended Dublin Core metadata standard to record entries or items into the database. I have used the References field to insert semantic linkages to Geonames codes. Special plugins provide functionalities to create structured vocabularies, to display places on maps and to create relationships between entries, based notably on standards defined by the Bibliographic Ontology (BIBO), the Friend of a Friend vocabulary (FOAF) or the Functional Requirements for Bibliographic Records (FRBR). It does not, however, create relationships between fields  orbetween  entities.

While Omeka’s front end offers generic templates and search outputs in acceptable format,  the power of Omeka come+s to its own with its Exhibit Builder, which is one of its most attractive features. It allows to easily create web pages of explanatory text with supporting images and other entries  by associating items in creative and historically meaningful ways; the visitor of the website may then navigate through a narrative. In this example I have used records regarding Gialova near Navarino to tell a story of successive productivity, war, desolation, disease and environmental management from the eighteenth century to the present..

Omeka is designed to serve individual researchers, research teams, teaching purposes, museums, libraries to publish their collections on the web, other scholarly collections and exhibitions but is unsuitable for  hierarchical structures, most notably for archives. Tellingly, for the Center’s Online Catalog the Rockefeller Archive Center uses a custom-made archival management system that complies with the Encoded Archival Description standard. It used, however, Omeka for a digital exhibition that celebrated the Rockefeller Foundation centenary with images and documents from its archive organised into themes that support a narrative.


The census project is entirely different. It involves a Medico-statistical survey of Greece compiled by state physicians between 1838 and 1840. It consists of hierarchies of tables of scientific and sociological data which are organised in a dated, complex and inconsistent structure that jars with current methods of data collection. It comprises spatial and prosopographical data, in other words it is a mine of historical information which the project wishes initially to record in digital format and then to explore and submit to research questions. Heurist, which stores the data in a MySQL database. seemed and proved to be a  tool suitable for the census project.

The survey tabulates by administrative department, municipality and community data pertaining to their physical geographic and climatological features, agricultural production, occupations and population data, schools,  prevailing diseases and their causes and the names of the medical personnel. This medical survey method was in tune with the medical-geographic approach to medicine of the day.

As is probably clear from the diagram of the database structure, all entities (or record types) are associated, one way or another, with entities place and/or person. This is achieved by field pointers that associate one or more fields of the source record to corresponding fields in the target record.  I have preserved the distinct tables of the original source as separate entities (or record types) and have also broken down entities with a large number of columns into separate entities. For instance the physicians, midwives and other medical personnel which were listed in the communities table were entered in a separate entity which consists of prosopographical and any career-related fields, for instance gender,  licensing dates and religion. I have also played with the idea of connecting the entries of the Landscapes database into the medicostatistical survey as a landscape record type with links to the Landscapes site.

As is suggested in the next slide, relationships between records may be defined by means of structured vocabularies that denote relationships,  in such a way as to construct hierarchical, kinship, temporal or any other imaginable relations, feudal, financial - whatever, reciprocal or not. One may even define the duration of a relationship, for instance the duration of an administrative dependence whenever administrative boundaries change. And, of course, it is easy to create semantic links to sources like Geonames. I use semantic links to Geonames owing to the high degree of granularity of topographical information of that database, certainly higher than the Getty Thesaurus of Geographic Names and, for the settlement patterns of the modern era, higher also than the Pleiades. Although geonames has accuracy issues, it provides information on altitude, which is important.


The following slide illustrates the result of a simple single-term search, such as the search for “Skardamoula” for instance; the output reveals the complex position of that particular place in a constellation of records, which include other communities and persons.  Another search method, namely faceted searches, facilitates complex searches through fields in linked entities and retrieves a network of connections,  like the one illustrated in the following slide which returns places suffering from intermittent fevers which lay below 500 meters from sea level. [Intermittent fevers was a nosological entity that broadly corresponds to malaria.] Maps are not limited to Google or Openstreet layers; one can incorporate shapefiles, kml format files and georeferenced images with little trouble.


To end my presentation: Eventually, completing the database project with Heurist will require the expertise of an IT professional to design the front-end website that will give outside access to the data of the survey. I have just managed to produce very rudimentary webpages like the one listing the names of midwives.
slides

2 comments:

  1. Paper presented at the 2nd CAA-GR confercence, Athens, 20-21 December 2016

    ReplyDelete

Note: Only a member of this blog may post a comment.