In the midst of all the hullabaloo about Ron Paul's decades-old racists newsletters, and his denial that he "never read that stuff," I ran across an interesting attempt to adding some data to the discussion. Filed under "pointless data exercises" and "politics," blogger Peter Larson has used text analysis to compare his blog, Ron Paul's recent speeches, and the original newsletters. He calls his results a smoking gun, with a question mark tacked on, and argues that Ron Paul wrote most of the original newsletters.
Before I go on, let me say that this is a really cool application. Instead of the he-said-she-said debate that's running in the media, this piece brings some actual data to bear on the conversation. Bravo!
That said, now I'm going to harp on statistics and inference. The problem with Larson's analysis is that he never addresses the question, "If Ron Paul didn't write the newsletters, who did?" Without answering that question, and putting some probabilities behind it, it's going to be very hard for text analysis to prove the issue one way or another. (Larson admits this on his blog.) Right now, his analysis proves that between himself and Ron Paul, Paul is much more likely to have written the letters. Not exactly a smoking gun.
That said, it's interesting data. For the record, I'm mostly convinced. In my mind, the statistics are flawed, but they still lend some weight against Paul. Oddly enough, Larson's finding that several of the letters were probably *not* written by Ron Paul was particularly persuasive. It feels human and messy, the way I would expect this kind of thing to be.
Final thoughts, mainly intended for my statistically minded friends: yes, this is an unabashedly Bayesian perspective. I'm demanding priors, which probably can't be specified to anyone's satisfaction.
IMHO, the frequentist approach has even deeper problems. From a frequentist perspective (which is where Larson's original, and deeply confusing p-values come from.) we use Paul's recent speeches and writings to estimate some parameters of his current text-generating process. We then compare the newsletters to that process and estimate the probability that the older text was generated by the same process.
Problem: we *know* the old text was not generated by the same process. It was written (allegedly) by a younger Ron Paul, on different topics, speaking into a different political climate. Without a broader framework, it's impossible to determine whether the differences are important. The Bayesian approach provides a direct way of assessing that framework. The frequentist approach doesn't -- at least not that I can see, without jumping through a lot of hoops -- and in the meantime, it obscures the test that's actually being conducted.
Thoughts on computation, social science, and lifehacking
from an up-and-coming data scientist.
Friday, December 30, 2011
Thursday, December 29, 2011
Job posting on Polmeth: Research Officer in Quantitative Text Analysis
This came across the polmeth mailing list a couple hours ago. Looks interesting, but the pay isn't exactly competitive for candidates with "a postgraduate degree in computer science, computational linguistics, or possibly a cognate social science discipline."
Job opening:
Research Officer in Quantitative Text Analysis
Duration: 24 months
Start Date: 1 March 2012 or as soon as possible thereafter
Salary: £31,998 – £38,737 p.a. incl.
Applications are invited for the post of Research Officer, to assist with a principal research officer for the European Research Council funded grant Quantitative Text Analysis for the Social Sciences (QUANTESS), working with Professor Kenneth Benoit (PI).
The research officer’s general duties and responsibilities will be to work with text and the computer organization, storage, processing, and analysis of text. These tasks will involve a combination of programming, database work, and statistical computation. The texts will be drawn from social, political, legal, and commercial examples, and will have a primarily social science focus. The research officer will be expected to work with existing tools for text analysis, actively participate in the development of new tools, and participate in the application of these tools to the social scientific analysis of textual data.
The successful applicant will be expected to possess advanced skills and experience with computer programming, especially the ability use a language used in text processing such as Python; familiarity with SQL; and experience with the R statistical package or the ability to learn R.
The successful candidate should have a postgraduate degree in computer science, computational linguistics, or possibly a cognate social science discipline and have an interest in text analysis and quantitative linguistics, have a knowledge of social science statistics and have worked in a research environment previously.
To apply for this post please go to http://www.lse.ac.uk/JobsatLSE and select “Visit the ONLINE RECRUITMENT SYSTEM web page”. If you have any queries about applying on the online system, please call 020 7955 6656 or email hr.jobs@lse.ac.uk quoting reference 1223243.
Closing date for receipt of applications is: 31 January 2012 by 23.59 (UK time).
Please access the attached hyperlink for an important electronic communications disclaimer: http://lse.ac.uk/emailDisclaimer
******************************
****************************
Political Methodology E-Mail List
Editors: Diana O'Brien <dzobrien@wustl.edu>
Jon C. Rogowski <rogowski.jon@gmail.com>
**********************************************************
Send messages to polmeth@artsci.wustl.edu
To join the list, cancel your subscription, or modify
your subscription settings visit:
http://polmeth.wustl.edu/polmeth.php
**********************************************************
Software design for analytics: A manifesto in alpha
If you take the lean startup ideas of quick iteration, learning, and hypothesis checking seriously, then it makes sense to build your software in a way that lends itself to doing analytics. Lately, I've been doing a lot of both (software development and analytics), so I've been thinking about how to help them play nice together.
Seems to me that MVP/MVC (or even MVVM, if you're into that kind of thing) are good at the following:
Having worked with (and built) several different systems at this point, I've realized that some designs make analytics easier, and some make them much, much harder. And nothing in general-purpose guidelines for good software design guarantees good design for analytics.
Since so much of what I do is analytics, I'd like to ferret out some best practices for that kind of development. I don't have any settled ideas yet, but I thought I'd put some observations on paper.
Some general ideas:
I'll close with questions. What else belongs in this list? Are there other people who are thinking about similar issues? What process and technical solutions could help? NoSQL and functional programming come to mind, but I haven't thought through the details.
Seems to me that MVP/MVC (or even MVVM, if you're into that kind of thing) are good at the following:
- User experience
- Database performance
- Debugging
Having worked with (and built) several different systems at this point, I've realized that some designs make analytics easier, and some make them much, much harder. And nothing in general-purpose guidelines for good software design guarantees good design for analytics.
Since so much of what I do is analytics, I'd like to ferret out some best practices for that kind of development. I don't have any settled ideas yet, but I thought I'd put some observations on paper.
Some general ideas:
- Merges are a pain point. When doing analytics, I spend a large fraction of my time merging and converting data. Seems like there ought to be some good practices/tools to take away some of the pain.
- Visualization is also a pain point, but I'm less optimistic about fixing it. There's a lot of art to good visualization.
- Units of analysis might be a good place to begin/focus thinking. They tend to change less often than variables, and many design issues for archiving, merging, reporting, and hypothesis testing focus on units of analysis.
- The most important unit of analysis is probably the user, because most leap-of-faith assumptions center on users and markets, and because people are just plain complicated. In some situations (e.g. B2B), the unit of analysis might be a group or organization, but even then, users are going to play an important role.
- Make it easy to keep track of where the data come from! Any time you change the
- From a statistical perspective, we probably want to assume independence all over the place for simplicity -- but be aware that that's what we're doing! For instance, it might make sense to assume that sessions are independent, even though they're actually linked across users.
- User segmentation seems like an underexploited area. That is, most UI optimization is done using A-B testing, which optimizes for the "average user." But in many cases, it could be very useful to try to segment the population into sub-populations, and figure out how their needs are different. This won't work when we only have short interactions with anonymous users. But if we have some history or background data (e.g. FB graph info), if could be a very powerful tool.
- Corollary: grab user data whenever it's cheap.
I'll close with questions. What else belongs in this list? Are there other people who are thinking about similar issues? What process and technical solutions could help? NoSQL and functional programming come to mind, but I haven't thought through the details.
Subscribe to:
Posts (Atom)