Friday, 3 April 2015

Quant Peace Research Part 1: Whose truth/justice/reconciliation?


This is the first in a series of posts on quantitative peace research.  Relatively new and often relegated to the 'huh?' panel stream at conferences, quant research in peace and conflict studies has been gaining ground in the past decade for several reasons.  Besides imbuing the word peace with less fuzzy/idealistic qualities thereby making it easier to build policies, raise money, frame political action platforms, and fuel solutions to conflict (Although it is the source of conflict with qualitative purists.  I won't hide that I am pro mixed methods... a topic for another post), the most promising reason quant research is having an impact in the field is that studies are being conducted by researchers with some surprising backgrounds beyond the usual polisci/sociology/social science disciplines.  These individuals bring with them alien concepts from neuroscience and even field experience from actual military combat that make for a truly multidimensional approach.

Whose Truth/Justice/Reconciliation? is a short version of an upcoming talk at The 8th Human Welfare in Conflict Conference. 

While presenting a seminar on the findings from my doctoral research (looking at narrative distortions produced via mobile ICT applications by bilingual participants), someone made an interesting comment about the applicability of my research, and methodology in particular, to transitional justice contexts.  This paper is an opportunity to develop that idea further.

In very generalized terms, post-conflict peace processes known often as ‘truth and reconciliation’ programs have been criticized for imposing a framework of justice that is culturally mismatched to the participating population’s concepts of justice. (see, for example, Avruch, 2010)  In this paper, I focused on concepts of culpability and agency by providing a nuanced quantitative measure to distinguish culturally rooted concepts of justice, responsibility, and agency. (Boroditsky, 2010; Costa et al., 2014)  The novel methodology adapted from cognitive linguistics combined quantitative measures of conceptual frames surrounding doubt, agency, and event structure to describe the concept of culpability.  Results have the potential to enhance dimensionality for articulating complex processes such as justice and  reconciliation as well as discussing the efficacy of such post-conflict programs.

I began by looking at the influence of ICT in conflict contexts because it is inescapable.  Due to the increasing use of information and communication technology (ICT) applications in the fields of peacebuilding and conflict resolution for gathering human rights abuse reports, election monitoring, polling, violence reporting, and other conflict management data collection activities that inform policy-making and participatory governance, this research performed a bilingual experiment with methodology from cognitive linguistics in order to describe the problematic nature of the ICT used in the conflict management context. This study was the first to incorporate cognitive-level communication variations and preferences as design considerations in the context of conflict management. (a.k.a radical alien interdisciplinary research) 

I have written extensively in previous posts about what cognitive-level means, but a brief recap. Pulling in research from both cognitive psychology and linguistics that examines memory, thought patterns such as categorization, problem-solving, cause-effect relationships, concepts of time and space, and use of language, this methodology focused, in particular, on conceptualization.  Conceptualization often requires complex relational understandings of objects, persons, time, space, and events.  The type of concepts that interest me concern events such as those that might be reported during conflict such as violence at a polling station or other incidents recalled as narratives and collected with mobile ICT applications.  (These narratives, once aggregated, become data for policy makers indicating hotspots of violence, political unrest, economic need, or even health crises.)  In order to observe concepts (because I can't see thoughts), I observed 'conceptual frames.'  A frame is something that computer scientists refer to, cartoonists refer to, cognitive psychologists refer to.  It is a fragment like a subject or a predicate, a basic unit of cognitive capacity that describes a perception such as he vs. they or something falling vs. something rising.    

Building from earlier work I had done which examined the consequences of a language barrier for ICT in crisis contexts (post-earthquake Haiti 2010, Libya and Egypt 2011, and Somalia 2011/12) which asserted that:
This flawed application prevented the original contributors from interacting with the information directly related to their own life-threatening situation, and the information it amassed formed an unsound basis for decision-making by international actors…. (Sutherlin, 2013, p.1)
my doctoral research (as well as this new paper) pursued the idea that the conceptual structure underlying language—the ‘organizational logic’ that occurs at the cognitive or thought-level—remained problematic for participation with ICT tools and the power they can leverage for policy-making for use by local actors.  In order to investigate conceptual structures, this research adapted experiments from cognitive linguistics that provided a quantitative means to assess the communication of concepts.  

In the northern region of Uganda, Gulu district, Gulu town, 29 bilingual Acholi-English participants completed a three-stage experiment. (I know it doesn't sound like a lot but that's an average number of participants for this type of bilingual study.)  Participants viewed a YouTube video depicting a chaotic street brawl, and were then asked to describe what they had seen in three distinct narrative forms: oral Acholi, written Acholi on a mobile device, and oral English.  By comparing narrative construction and identifying concepts unique to certain narratives, the experiment looked at the level of thought before language, the cognitive level, and thus followed in the footsteps of earlier research in the field of cognitive linguistics that examined how concepts from one language can be observed to transfer into another.  

**A quick note about language/culture/cognition: because this was a bilingual experiment and the data was in the form of 'language' but the variables under investigation were cultural and cognitive variation, the two comparison languages used in the experiment should be considered exemplars of cultures with certain characteristics that have a high cognitive impact such as orality or how categories are used. The characteristics which differ form a really long list, but part of the reason these two languages are compared/contrasted is their linguistic and cultural distance to one another which brings the issues under investigation into relief.  (Linguistic distance was  proposed by Greenberg in 1956 and extended by Lieberson in 1964 and even has a Wikipedia entry so it has got to be pretty well established. It quantifies how different dialects and languages such as German and Dutch vary from one another.  Cultural distance is adapted from this idea.  So English represents the culture that produced the ICT application and Acholi represents the culture using the application for data collection/aggregation/policymaking in conflict contexts.)  To summarize, it's not an experiment about English vs. every other single language or any specific language at all; it's about variations in underlying thinking (preceding or accompanying language).  Because we can 'see' thinking, the experiment observes language production and makes inferences about cognition and the culture that influenced it.

During the analysis, I looked for evidence of English to transfer into Acholi due to the presence of ICT.  For example, in the ICT recall stage, although participants were reading in Acholi and writing in Acholi, the logic of the ICT application which had been designed (as nearly all software has been) with the logic of English in its core would trigger English concepts in participants bilingual brains.  Concepts from English would transfer into their Acholi narratives that would not normally appear in an Acholi narrative.  My hypothesis was that, in essence, the ICT format would prescribe the participants' narratives in a way that was not natural to Acholi; there would be distortions or dissonance.  In 3/4 of the cases this was true.  There was a narrative shift and not simply one attributable to speaking vs. writing because I was looking at specific schemata and narrative structure. (From speaking to writing you might change how you describe something, but you don't change the story.)   

Before conducting the experiment, I spent three months in the field.  I did intensive language immersion.  I had discussions with local university professors, hunted for literature to review (anthropology, literature, poetry, linguistics, narrative studies, psychology).  All so I could identify specific cultural schema and narrative patterns. Schemata (sing. schema) are cognitive shortcuts that our brains use to make sense of the immense amount of sensory information we take in.  They are made up of conceptual frames.  For example, you can recognize a dog in a fraction of a second out of the corner of your eye because it fits the model/shortcut/set of conceptual frames for that animal.  We rely on schemata in order to be more efficient with our mental energy as well as to make sense of unusual or new situations by slotting what we see/hear/etc., onto the scaffolding of existing a schema and proceeding with a 'best fit' guess.  By focusing on schemata, this connected the experimental results to culturally formed concepts and the level of thought rather than a discourse analysis on language.  Schemata are culturally informed in this way-- you are probably familiar with the adage, 'When you hear hoof beats think horses, not zebras.'  Does everyone everywhere think horses?  It may depend on place/culture.   The video prompt for the experiment was chaotic and shared some familiar characteristics (because it was a street scene in Nigeria and the market stalls and taxi stand looked similar to Uganda as well as YouTube having made Nigerian videos popular viewing across the continent); however, the unfamiliar language in the video and, again, the chaotic scene, made it likely that it would trigger in participants the reliance on their culturally learned schemata.  That was the idea anyway.

Among the key findings, the concepts of culpability (who was guilty) and agency (who was involved) emerged as unique between what was described via ICT and orally in Acholi.  Crucially, several participants claimed that one specific individual was to blame for the incident in their ICT recall while they had only described a group having possibly been involved in something during their initial Acholi oral recall.  In addition, several participants changed the very nature of the event between these two recalls.  If we imagine these reports as part of a police investigation, the initial set of oral reports seems to indicate no action is needed while the ICT reports point the finger at one man.  Troubling to say the least.

In conclusion, if cultural constructs such as justice, culpability, and agency are both consciously and unconsciously programmed into technology, then the ICT application is putting limitations on the narrative, perhaps even prescribing conceptual elements of narrative for something as vital and nuanced as justice.  If we imagine a field poll being taken about what form transitional justice should take, if technology is involved, even in the aggregation of narratives later, this could radically alter the results by altering authorship/intentionality/voice/participation.  In addition to this practical impact, the methodology I used (with or without the mediating factor of technology) could offer a deeper understanding of the core conceptualization of justice within a society by being able to break the concept down at a cognitive level.  Subsequent posts will continue to look at each of these concepts in more depth (culpability and agency) as well as build on comments/reactions to the paper.

Tuesday, 24 March 2015

Hibernation is over

There has been a period of hibernation on this blog due to the completion and defense of my PhD last fall/winter.  And while this platform was primarily created as a space to explore topics related to my doctoral research, in the immediate future, I do plan to continue to use it to informally develop ideas as I speak and publish in the areas of peace & conflict, (cognitive) linguistics, information science, and their intersections with culture and politics.

Coming soon... a focus on quantitative peace research beginning with some thoughts about the communication strategies of ISIS from a cognitive linguistics perspective with comments on the report from the Brookings Institute integrated with some other projects such as the peace and terrorism indices.

Also coming up... thoughts on paper to be delivered at the 8th annual Conference on Human Welfare in Conflict at Oxford's Green Templeton College, titled: Cultural Concepts of Culpability: the role of ICT in post-conflict transitional justice. 

And in the meantime, I just came across this painting in the MFA Houston, and it resonated for 2 reasons.  First, it has an obvious conflict resolution context.  But second, it reminded of the optical illusions that integrate two images (the old/young lady) so that viewers see different images depending on their perspectives.  In this painting, there is no optical illusion at play, but it is simultaneously urgent and frenzied at one end and deliberate and calm at the other.  It depends on the viewer's perspective which aspect controls the scene's narrative.  And so it goes in conflict res....
During my phd, I did learn the secret to keeping elephants out... where does that go on my cv?


Monday, 20 October 2014

Think Outside the Lab


This takes me back full circle to the initial post with which I launched this blog—my communicative purpose and challenge was to work within an impossible sort of Venn Diagram of three disciplines that never seem to collaborate: Computer science (such as developers, information scientists, and engineers), linguistics (including cognitive linguistics, but really all language-oriented studies), and finally political science (policy-oriented and often practitioners such as human rights activists or crisis response managers).  Pairs of these disciplines can be found teaming up, but an effort combining the insights of all three is, sadly, very rare indeed. 

A recent project out of MIT from Berzak, Reichart, and Katz pursued the hypothesis that structural features of a speaker's first language will transfer into written English as a Second Language (ESL) and can be used to predict the first language of the speaker.  I believe their work does not go far enough in two respects.  First, in terms of sampling.  Second, in terms of considering how parsing communication in this manner might be applied to software design.

Their paper addresses the sampling problem as one of resources.  They acknowledge that there are over 7000 languages, and there exists a written corpus (their data pool) from only a relative few.  My critique, however, is that they consider all languages members of the same sample set for their experiment.  Katz explains in an interview that he was drawn to investigate and algorithmically describe ‘mistakes’ made by Russian speakers in English.  These mistakes are called linguistic transfer because an element from the first language is transferred into the second.  (Reverse transfer can happen as well when a new language affects the first language.)  Linguistic transfer can come in several forms: phonological/orthographic (mistakes due to sound or spelling), lexical/semantic ('false friends'), morphological/syntactic (grammar mistakes), sociological/discursive (such as appropriateness or formality), and conceptual (categories, inferences, event elements, concepts generally).  If Katz’s group had differentiated between types of mistakes, they might have improved the rate of prediction success in their results. Also, it is unclear if their model was able to incorporate more complex types of transfer such as discursive or conceptual.  One reason they may not have seen the need to differentiate by type was their limited sample. 

Most linguistics studies (or research asserting multi-lingual or multi-cultural value) that purport to incorporate a broad range of languages, in fact, only draw from a small number of closely related languages that don’t possess particularly profound differences in conceptual organization of information. That means if there were instances of conceptual transfer, they would be rare or at least difficult to detect.  (Most studies look at Indo-European languages, plus perhaps Russian, Hebrew, Korean, or Japanese to appear to have real diversity.)  Among the nearly 7000 languages, there are only 100 or so that have a literature; it is this group of languages that are most frequently studied. These languages are, therefore, ones which have a strong history and preference for writing (called chirographic), and this mode of communication has had an effect on many cognitive processes within the populations that speak these languages.  The rest of the 7000 are predominantly oral, and there are very rarely oral languages represented in the sample sets (of any study).  Orality is not be confused with literacy; it is a preference for communication and most speakers of predominantly oral languages also speak and operate in chirographic languages as well.  The impact on cognitive processes such as categorization, problem solving, ordering for memory, imagination, memory recall, etc, is connected to a need to rely on sound and associated mneumonics for information organization.  If you cannot write something down, this changes your strategy for remembering something or for working through a problem or any number of other cognitive processes.  The linguistics studies that fail to represent a member from this set of predominantly oral languages make an egregious sampling error which leads to false conclusions about universal or easily modeled qualities of communication.  Orality is a profound variable in terms of its effect on cognitive processes. That is why investigating and describing communication at a conceptual level and drawing from languages much more distant to the typical baseline of English would yield some surprising results.

The second problem with the MIT study is one of anticipating a use for the findings.  Quoting from the press release about their work:
"These [linguistic] features that our system is learning are of course, on one hand, of nice theoretical interest for linguists,” says Boris Katz, a principal research scientist at MIT’s Computer Science and Artificial Intelligence Laboratory and one of the leaders of the new work. “But on the other, they’re beginning to be used more and more often in applications. Everybody’s very interested in building computational tools for world languages, but in order to build them, you need these features. So we may be able to do much more than just learn linguistic features. … These features could be extremely valuable for creating better parsers, better speech-recognizers, better natural-language translators, and so forth." (L. Hardesty for MIT news office 2014)

Yes, so true.  However, it's not theoretical at all, nor is it simply the folly of linguists to pursue communication variation at a conceptual level.  Using conceptual frames (Minsky, 1974) has already been proven to be an effective method in improving the search capability in map tools by Chengyang et al. (2009)  who shifted a map search tool to operate from conceptual frames rather than conventional English search terms.

Katz and his lab at MIT are credited with the work the led to Siri, and this new study could be applied to machine language tools so that patterns of mistakes become predictable and thus correctable.  It could also be added to text scanning tools to detect the first language of non-native English authors on the web thus adding to the mass surveillance toolkit.

I think a lot more could be done (but hasn't) with this methodology in terms of looking at how an oral language's conceptual frames could be described and then used to calibrate a more responsive information and communication application.  I used very similar methodology to the MIT researchers in my experiment looking at reverse linguistic transfer last year (that I have been charting on this blog).  I compared sets of bilingual narratives and looked for patterns of 'mistakes,' but I was interested in what these mistakes could tell us about the communication needs of the users (mobile technology users in rapidly growing markets like Africa, South East Asia, or South America).  My hypothesis was that the structures of their first language were being  distorted, being converted into 'mistakes,' in order to fit a prescribed (foreign) conceptual structure of the software application.  What I found was much more complex than counting instances of mistakes.

What I observed, and quantified, was that when comparing the first language oral narrative to the first language narrative via mobile report (either as an SMS or as a smart phone app question series),  3/4 of the participants expressed dramatically different narratives when using the mobile report format than in their initial first language oral narratives.  That means that translating interfaces isn't sufficient to provide communication access.  There are underlying conceptual aspects to communication that have yet to be addressed and that are inherently cultural (currently mono-cultural).  Due to the complex nature of concepts such as justice, personhood, time, or place identifying and isolating instances of transfer was very challenging.  A summary of the results is forthcoming in 2 papers as well as my doctoral research, but the main conclusion I researched was conceptual-level parsing of communication should be integrated into design of communication and information management software with the integration of insights from oral languages.  Inclusion of this variable with indigenous software design will increase the ability of users from rapidly growing markets to participate with and leverage the information and communication technology in a manner which meets their needs.

this topic will be continued with highlights from forthcoming publications.

references:
Chengyang, Z., Yan, H., Rada, M. and Hector, C., 2009. A Natural Language Interface for Crime-Related Spatial Queries. In: Proceedings of IEEE Intelligence and Security Informatics, Dallas, TX, 2009. 

Jarvis, S. and Crossley, S. 2012. Approaching language transfer through text classification: Explorations in the detection-based approach. Multilingual Matters, volume 64.

Jarvis, S. and Pavlenko, A. 2007. Crosslinguistic Influence in Language and Cognition. London; New York: Routledge.

Tuesday, 30 September 2014

The Post-Terror Generation

When did generations stop being defined in terms of war?  In terms of the hardship that fueled dreams of a brighter future for the generation to follow?  The name of an age once served as a reference and reminder to ill-fated political policies, to the suffering wrought by our own hubris so that we might shield future generations from the mistakes that have cost us our chance.  When we examined a generation, their defining characteristic has most often been forged by war.

The Lost Generation who witnessed the hardship of WWI and the political forces that disappeared the golden age into the modern age.  The Victorian Age before it, synonymous with colonialist expansion and industrialization.  Generation X was the first generation to be defined without direct reference to a defining war or hardship because theirs was a generation without a cause, without fight or direction.  It was defined by its 'anti' characteristics, most notably apathy, in contrast to its parent generation, the post-WWII Baby Boomers.

There is now a post-terror generation.  Children born near the end of the 20th century and certainly after 9/11 who have never known a world that was not caught up in the War on Terror, who have not been consumed at every cultural level by a political policy structure reacting to the War on Terror.

The post-terror generation was the term used by Edward Snowden in an interview with Vanity Fair in April 2014  to describe the millennials, a group that polling research has noticed is a more optimistic, less polarized generation.  Trying not to repeat the behaviors they've witnessed in the generation that preceded them such as the US engaged in a multi-front war, a contentious, non-stop screaming match between political parties, and more and more leaders that find it easier to blow up a problem rather than reach out and build a solution (Generation Terror by Michael Scarfo). The post-terror generation will look past the bombastic pressures that rile or depress the rest of us as simply white noise. 

The War on Terror is a black hole of policies that have pulled us into war and obscured our focus on the environment, the banking collapse, poverty, education, healthcare..... creating an paralleled atmosphere of partisan rancor.  The post-terror generation had no other course but to rebel against this failed model.  Ironically, because they have always had the looming threat of terror, the fear does not govern their lives in the way the previous generation felt a tectonic shift had robbed them of an inalienable security.   

The evolution of current security policy is certainly more complex than this post delves into; here, I am interested in pondering how this young generation coming up behind me sees the world, and why we don't have cooler names for generations!  The alphabet system has run its course.  I vote for a return to connecting ourselves to a defining struggle.  The post-terror generation suits these kids.

Wednesday, 20 August 2014

The Slow Drip Invasion: use of ICT and UAV in weak states

If you imagine the challenges faced by local communities plagued by conflict and institutional instability, someplace for example like Somalia, a nation that has faced profound governance issues (sorry to pick on Somalia, but this post is about weak states), and you work in any part of the humanitarian or development organization network, it may seem perfectly reasonable to empower local communities and civil society groups to collaborate on Alternative Modes of Governance.  In this way, communities can see to their own basic needs.  Mobile technology has emerged as a resource that development experts are thinking creatively about in order tackle these types of issues.  The rapid influx of phones, more generally of information and communication technology (ICT) has presented new opportunities for governance according several influential tech architects.  In a new book, Bits and Atoms: Information and Communication Technology in Areas of Limited Statehood edited by Steven Livingston and Gregor Walter-Drop, contributing authors such as Patrick Meier the developer of Ushahidi (an SMS platform relied on by branches of the UN and US government) as well as development heavyweights such as Dr. Sharath Srinivasan, Director of the Centre of Governance and Human Rights (CGHR) at the University of Cambridge, explore ICT's viability as an Alternative Governance Modality.  They discuss the effects of ICT proliferation within the slums of Kenya and Russia as well as other areas that are considered to have limited or weak central governments.  

Alternative Governance Modality.  It sounds like a very neat solution.  Certainly seems like the best option in the slums of Kibera, Kenya.  If you break it down, it becomes decidedly less neat.  Alternative to what exactly?  If local community groups and NGO/civil society partnerships have been capacitated with ICT, where did that technology come from?  Where did the policy directives come from?  Is there a Dutch NGO or USAID project that has essentially invaded a small corner of some weak state that is unable to object, all via mobile device?  And where does the data go?  Who is it for?   
I am not advocating that communities should not organize or utilize whatever means they find in order to address the issues they face; however, the authors are not being entirely honest in their assessment of 'locally empowering' when they describe the use of ICT in these projects.  The tech tools employed are inherently external and foreign objects.  Indigenous ICT simply does not exist yet, and the means to develop it might, to incorporate the cultural nuance of information and communication preferences from Kibera or a Somali community, but these design techniques are not being used in the humanitarian ICT field.  The focus is mistakenly on simplicity, assuming that streamlining applications will overcome literacy issues or even culture barriers.  This approach compounds the problem.  What Western designers understand as the most logical, the most simple, the most intuitive, inherently expresses their conceptualization of how information should be organized and how it connects.  The ways in which information can be connected and organized at a conceptual level is by no means universal, particularly if you take into account differences between more predominately oral cultures.  (Check out some earlier posts on orality; it’s not the opposite of literacy, but a cultural communication preference with cognitive implications.) By concentrating or streamlining the design, it becomes extra-Western in its conceptualization and thus even more distorting to non-Western information and communication intentions.  The current interface and information design empowers Western users not the users described in these ‘weak state’ contexts.  By capacitating local groups with ICT, the authors are describing a situation for linking up new populations to a vast data network.  Connecting them as potential sources of information and points of leverage (perhaps for Western policy makers, perhaps for commercial enterprise).  ICT which captures information in ways useful to non-Western users, ICT that functions as a tool, a potential policy-making aid or technological advancement outside the Western concept has not been fully realized; therefore, what does exist is only useful or empowering to the group that designed it, that released it in the field, that is writing about its vast potential because that potential will be for them.  

The problem of data collection in weak states is brought into relief with the use of UAVs (drones). 
 The images they provide are meant to enable humanitarian crisis responders to more efficiently get know ‘the lay of the land’ when called to work.  Who could argue with technology that improves humanitarian missions and potentially saves lives?  This was the original purpose of ICTs like Ushahidi-- crisis response.  There is a jump in the script, a missing (or a few missing) steps that take us from designing a successful tool for crisis response in which information is marvelously organized and communications are streamlined during a short-term intensive mission by external actors to a stage where this technology is being used to govern long-term by indigenous populations.  These are two massively different tasks not to mention two different sets of users.  How did these parameters escape the designers?  How did we just slide into alternative governance modality from what was initially a cobbled together system to organize the hectic atmosphere of crisis response?  How did these projects go from responding to crisis (after the fact) to inserting into the fabric, the airspace, of weak states in an on-going capacity with the stated aim of preparedness?  The reason mapping projects seem to empower the technology providers more than the local population is that, for example, in a region where I recently did fieldwork in Northern Uganda the language spoken there has no word for map. This was true for the surrounding languages ranging into Ethiopia, Kenya, and South Sudan.  (read more about about it in Without a Map, and Like It's 1899).   The data collected with UAVs will arguably improve humanitarian missions.  But what else besides?  

There is something in scientific research called dual-use technology and for which certain safety protocols are developed.  A scientist may discover an amazing virus that can be harnessed to cure cancer, but it could also be released as a weapon—there are two uses.  So far, ICT applications and other digital technology are not treated as creations with this same bipolar potency.  There is little ethical debate about the long-term implications or context of use.  There is certainly no ethical training for designers or engineers.  I have been pleasantly surprised by a few technology journal editors that encourage ethically driven arguments, but I think it would be terrific if there were more voices in the field taking the idea of 'empowerment,' a bit further, that is to say really delving into how this power comes about and from where/to whom.   To put it another way, these tools will only become empowering for more people, become better tools if designers are driven to improve them—this is an area where there is huge opportunity to develop new tools for new groups of users (staggering large groups of new users) that approach local governance or any number of issues from a non-Western conceptualization.  A total departure from the humanitarian crisis responder-user and task and an embrace of the indigenous user and his/her information and communication preferences should lead to a much more successful tool.  This culturally based ICT development is certainly on the horizon.  I can't wait to see (or hear) it in action.

Friday, 25 July 2014

Dark Cloud over Academic Freedom

The US supreme court's ruling upholding the subpoena issued on the basis of the Mutual Legal Assistance Treaty (MLAT) between the US and the UK makes researchers vulnerable to the same reprisals and targeting as informants and spies.  The work of researchers, the interviews they collect, the analysis they provide, can be invaluable in conflicts and this ruling changes them from a resource for peace to a tool for destruction.

What was the case that tipped the balance?  A murder in Belfast over 30 years ago allegedly committed by Gerry Adams.  (Prof. Robert White of Indiana University's sociology dept gives a nice timeline of The TroublesThe News Letter, The Pride of Northern Ireland gives a timeline of the events of the case; and Boston College gives a timeline of legal proceedings)  The reason a murder case in Northern Ireland touches reseach in the US is the MLATreaty... there were documents collected by researchers at Boston College that were subpoenaed as evidence, and similar the conundrum courts face with journalists when they know details of a crime, the court in Northern Ireland felt there was important evidence in the interviews that could be shared via this treaty.  The Belfast Project documented many hours of interviews with participants on both sides of the conflict on the strict written understanding of confidentiality until their death unless otherwise granted.  This kind of precaution was and still is felt to be necessary to protect interviewee's lives and those of their families.  In fact, this case is about a kidnapping and murder of a woman by the IRA.

Now there is ample meat on this bone of contention for legal and social science scholars.  First, should we consider the researchers at Boston College researchers or were they journalists or even IRA affliated persons (which gave them trust to interview those communities) with no academic credentials?  Does their categorization even matter when the real issue is the breach of confidentiality of their sources?  The breach of trust, a foundation in collecting interview-based research.

Another issue, not emphasized during any of the court hearings, was that after 1972, anyone arrested and imprisoned in Northern Ireland was considered a participant in the conflict rather than a criminal.  There was a conceptual and legal change in standing for crimes committed thereafter as being part of a larger battle.  As I understand it (from speaking to experts in this area), if this murder charge had been brought in 1972 and Adams had been arrested then, it would not have been a murder charge but rather a political one (called Special Catergory Status).   And this is a key distinction both during and post-conflict for rebuilding because, to take a different example like Egypt after the revolution in the streets in 2011, would it be helpful to go back and prosecute every person they could find for breaking a window for vandalism and every person for assault and battery for protest related violence?  (Certainly some post-conflict resolutions chose to pursue justice for key leaders such as through the ICC, but this is not always the case.)  At some point, most post-conflict societies decide to draw a line of forgiveness (such as truth and reconcilation) in order to move forward.  And in no way is the forgive/move on method easy, it is simply that there is something unusual about a murder case in these circumstances.

The researchers ceded their interview data to the Boston College Library where anonymized data could by used by other researchers.  The data was held by a third party not unlike how we use cloud data storage or email or other digital storage resources to facilitate data collection and security.  Ultimately, it became the university's decision to comply with the subpoena not the researchers themselves because they had given up the data.  How we store our data, who controlls it, who has access to it is ever more important with this implications of this ruling.

As argued in the Massachusetts ACLU's amicus brief, described here by their executive director Carol Rose, “It is alarming that the trial court opinion suggests that the Constitution surrenders US citizens to foreign powers with fewer safeguards than are afforded to citizens subpoenaed by domestic law enforcement agencies.  If the government has its way, it would straightjacket judicial review of investigations and prosecutions by any foreign country party to this treaty, including Russia and China.”

The examples given in the amicus brief illustrate how information sharing (or not sharing) was a factor in recent legal actions in countries subject to the MLAT:
The prosecution of Nobel Prize winner Liu Xiaobo by the Chinese government for, “inciting subversion of state power.”
The recent arrest and prosecutions of non-govermental organizations, including civil rights groups, by the Egyptian government.
The sex discrimination case recently dismissed by a Russian judge who stated that, “If we had no sexual harassment we would have no children.”

This begs the questions, why did the US Supreme Court grant this subpoena request now?  The support and 'special relationship' between the US and the UK has developed a unique flavor as a result of the war on terror.  A kind of complicity.  Despite pressure from senators and then Sec. of State Clinton, as being a politically destabilizing move, this ruling opens the door wider for government to pressure researchers for data.  For me and my colleagues, I can only imagine the consequences.  We are the ones hiking into the hills to ask former child soldiers about their experience, to ask suspected taliban about their motivations, to ask corrupt drug enforcement police about their allegiances, what could possibly go wrong for us or for the people we interview if we are are no longer seen as purely academic researchers?

And what of the critics who say that social science provides no concrete results towards solving war and conflict?  Just because the effects of research informing policy-making are too complex to throw up on a powerpoint slide does not mean they do not exist.  The knowledge gained by investigating the nature of conflict, its intricacies and ramifications, its participants and their motivations, this certainly leads to better planning for preventing conflict and better policy-making when embroiled the unstoppable ones.  What is the alternative?  Not understanding the nature of the thing and making guesses about policies for troops and sanctions and alliances in the dark?

Finally, the amicus brief written by group of concerned social scientists does a wonderful job of outlining several key reasons why this ruling was aggregous and should be added as a point of review on the ethics panels for all researchers in order to understand how their data will be protected at their institutions.  In fact, if you've never read a legal brief (or tried it and hated it), this is the one for you.  It tells a story, makes a compelling argument, and stays well clear of jargon and things like, 'pursuant to code 3.1.c.-3. blah blah.'  Enjoy.




 




Friday, 25 April 2014

The Blind Spot for Big Data

The New York Times has been doing a series of pieces on the uses and limitations of Big Data.   While I do not specifically focus on big data, I look at some of the ways we collect it; therefore, I am interested in the downstream implications once it's aggregated.  How could small distortions at the scale I study become much larger? 

Since I look at conflict, the piece by Somini Sengupta, 'Spreadsheets and Global Mayhem' certainly caught my eye.  The title for the opinion piece about all the ways we are trying to mine data for conflict prevention matches the term 'spreadsheets,' a feeble and not very advanced technology for organizing stuff, against the description 'global mayhem' (for me it evokes Microsoft Excel battling the Palestinian-Israeli conflict.)  The title conveyed the incongruence of strategies centered on big data.   Collecting information, aggregating it, that isn't enough.  The sheer weight of it, the potential feels powerful.  Surely, answers must be in there somewhere?  But finding patterns, asking the right questions, creating really good models with the complex information such as communications data (much of it translated)... that's a long ways off.  We don't really know what to do with what we have, and we don't really know what the answers mean from the models we build.  That's where I think we are.  Most marketing firms vehemently disagree. (Sentiment analysis).  And certainly the types of conflict prediction machines Sengupta references such as the GDELT Project and the University of Sydney's Atrocity Forecasting  believe fortune telling is within our digital grasp.

Another piece by Gary Marcus and Ernest Davis 'Eight (No, Nine!) Problems With Big Data' addresses some of these issues including translation.  They remind the reader of how often the data collected has been 'washed' or 'homogenized' with translation such as the ubiquitous Google Translate.  The original data may appear several times over in new forms because of this tool.  And there is a growing industry of writing about flaws with big data.  The debate has made many who work within the field weary or intensely frustrated because the debate is fueled largely by popular misunderstandings of a very complex undertaking.

From my perspective, there remains a giant blind spot, what I call the invisible variable of culture. Most acutely, this involves the languages now coming online, the languages spoken in regions experiencing a tech boom.  Individuals in these areas must either either participate online and with mobile communication technology in a European language or muddle through a transliteration of their own local language which will not be part of this Big Data mining.  My research looks at the distortions in the narratives they produce in both instances.  The distortion over computer mediated communication such as SMS or smart phone apps which compartmentalize narrative, is a problem about how we organize what we want to say before we say it.  This pre-language process varies by culture and structures how we connect information such as sensory perception.  At the moment, our technology primarily reflects one culture's notion of how to connect information, how to organize it conceptually.  This has implications both in how information technology collects data and how questions about that data are posed and understood.

What if other cultures have a fantastically different concept of organizing information?  How do you know the data you've collected means what you think it means?

[math example: your math is base 10... but other groups might use base 12 or base 2, etc... so when you see their numbers and analyze them with your base 10... they make sense to you but don't mean what they meant originally.]

We haven't cracked the code yet of how to incorporate a variable like culture into software applications.  It's more than translation.  It's not as easy as word replacement.  It's deeper than that.  It's context.  It's at the level of concepts and categories.  The way we see things before we use language.  That's not to say we can't unravel these things with algorithms... but those are often based on (even unconsciously) our understanding of communication.  And there is massively insufficient research on most languages out there.   If there are around 6800 languages, Evans and Levinson (2009) figure that:
Less than 10% of these languages have decent descriptions (full grammars and dictionaries). Consequently, nearly all generalizations about what is possible in human languages are based on a maximal 500 language sample (in practice, usually much smaller – Greenberg’s famous universals of language were based on 30), and almost every new language description still guarantees substantial surprises.
And the languages within the tech boom regions such as Africa and Southeast Asia are certainly part of the knowledge void.  We aren't prepared to collect this data yet.  The data we do collect are basically shoehorned into a format meant for English and for western concepts (like our notions of cause and effect or even time).  Data from these language groups including usage patterns, such as the flu or pregnancy predictor algorithms we've read about, won't be any good without further cultural adaptation.  And when it comes to crunching the data, we have a lot to learn about asking context specific questions and understanding the data from a non-western framework. (My own research results have shown me it's the difference between thinking you've identified a victim or a villain.)

While not widely understood yet, these cultural differences in the Big Data story are a dazzling challenge to consider.


The Global Database of Events, Language, and Tone (GDELT) is an initiative to construct a catalog of human societal-scale behavior and beliefs across all countries of the world, connecting every person, organization, location, count, theme, news source, and event across the planet into a single massive network that captures what's happening around the world, what its context is and who's involved, and how the world is feeling about it, every single day. - See more at: http://gdeltproject.org/index.html#sthash.Tl6hO993.dpuf
The Global Database of Events, Language, and Tone (GDELT) is an initiative to construct a catalog of human societal-scale behavior and beliefs across all countries of the world, connecting every person, organization, location, count, theme, news source, and event across the planet into a single massive network that captures what's happening around the world, what its context is and who's involved, and how the world is feeling about it, every single day. - See more at: http://gdeltproject.org/index.html#sthash.Tl6hO993.dpuf
The Global Database of Events, Language, and Tone (GDELT) is an initiative to construct a catalog of human societal-scale behavior and beliefs across all countries of the world, connecting every person, organization, location, count, theme, news source, and event across the planet into a single massive network that captures what's happening around the world, what its context is and who's involved, and how the world is feeling about it, every single day. - See more at: http://gdeltproject.org/index.html#sthash.Tl6hO993.dpuf
The Global Database of Events, Language, and Tone (GDELT) is an initiative to construct a catalog of human societal-scale behavior and beliefs across all countries of the world, connecting every person, organization, location, count, theme, news source, and event across the planet into a single massive network that captures what's happening around the world, what its context is and who's involved, and how the world is feeling about it, every single day. - See more at: http://gdeltproject.org/index.html#sthash.Tl6hO993.dpuf