Showing posts with label computational social science. Show all posts
Showing posts with label computational social science. Show all posts

Thursday, September 26, 2013

The Predictionary



See below: the methodology alone is a complete mindfuck, and should provide some sense of the nature of scientific studies in the age of big data.

Get ready...

Researchers say readers' identities can reveal much about content of articles
Aug 12, 2013
modified article:

Articles that people share on social networks can reveal a lot about those readers, but a  new study reverses the proposition: What can be learned about an article from the attributes of its readers?

To find out, the CMU researchers, along with colleagues at the University of Washington, analyzed almost 3 million news articles and the public profiles of the people who shared those articles on Twitter.

This enabled them to generate a few thousand "badges" that characterized the content of the shared news articles and also could be used to analyze any subsequent article, including those that had never been shared or even read.

In order to train their model, the team began by looking at three months of tweets—from September of 2010, 2011 and 2012—and selecting those that included links to mainstream news articles and came from a user who had filled out a Twitter profile.

[collect major news outlet's articles that have been tweeted about] 

Each news article was then downloaded and the most meaningful, unique words were extracted, creating a "bag of words" for each article; similar to a visual word cloud, these bags give greater weight to more important words. Likewise, from each user's Twitter profile, a set of descriptive words, or badges, was extracted.

[each article, as well as each twitter-user-profile, gets a weighted wordcloud]

By comparing the bags of words with badges from the people who shared the articles, the researchers were able to create a dictionary that associated each badge with its characteristic words. For example, people who self-identify with the music badge in their profiles are likely to share articles with words such as "band," "album" and "song." Different dictionaries were created for each year to compensate for interests or topics that change over time. These dictionaries were then used to encode new articles, leading to a document representation based on attributes of potential readers.

[the article wordclouds and the user-profile wordclouds of the users who tweeted those articles are cross-correlated to create a "dictionary", or rather a "predictionary", if you will, that predicts who will share what, or what will be shared by whom]

Case Study Example:
New York Times columnist Maureen Dowd had readers who tended to be progressive. This association was notable because Dowd never explicitly uses the word "progressive" in the articles analyzed by the researchers. Rather, the algorithm detected that the words Dowd uses in these articles correspond to the type of content self-described progressives tend to share on Twitter.
-Carnegie Mellon University


Thursday, March 7, 2013

Seeker to Uploader Ratios and Botsites


Sandra -psychosandra- Holmbom

It's tough sometimes to transpose tabs for the uke, believe it or not the majority of uke tabs for certain tunes are software-generated and often wrong. Something about the different ratios of uke-playing uploaders to uke-learning tab-seekers.

Is it because advertisers are more likely to generate botsites for services with higher seeker-to-uploader ratios?



I play the guitar. I also play the ukulele. I don't have much of an idea what I'm doing when I'm trying to figure out how to play a song that I like, as I've always been able to reference a guitar magazine, or because I've always tried to play pretty easy songs with simple chords.

Nowadays, I like to play the uke, and I like to play 7th chords, 9th chords, etc. When I want to learn a song on my uke, I look for online tabs, as I'm sure everyone does. In so doing, I've found that, for certain kinds of songs - typically older songs, like prior to the rise in popularity of the uke (~2010's) - the tabs that can be sourced-up are way off, if not totally wrong. What I usually end up doing, is to look for guitar tabs, and transpose them (which isn't very easy for someone like myself who never really learned music properly). Luckily, I learned some basic music theory in my initial years of guitar playing, and it's been relatively easy to augment that learning, at least enough to meet my needs of detecting the accuracy of my own transpositions.

*The ukulele, though it is a stringed instrument and it looks exactly like a small guitar, requires a different tablature than guitar. The strings are GCEA, not EADGBE, which makes the fingerings of the chords completely different.

Ukulele Makers, fiddle and uke playing robot

There's something about the guitar that makes it easy to seem like you know what you're doing, even if you don't. And the same goes for the uke, even moreso (simply because it has less strings). The uke rose to popularity only recently, and therefore after the rise of the internet and the death of hard-won knowledge (if I had the internet of today when I was learning guitar twenty years ago, I would never have tried to figure out songs by myself, and I would have understood much less of music theory because of it.) Players learning the uke today (and looking for or potentially uploading tabs), then would be less likely to have to learn as much to get by as a guitar-player of twenty years ago.

So the ukulele is somewhat easier to play due to both it's lesser number of strings, but also because of the means with which to learn new songs on it (which expands playing-availability to a wider audience of non-musically-trained players).

Finally, the guitar is a more widespread instrument than the uke, which brings with it more players who might know what they're doing, and tabulate and upload songs for others to learn from.

I speculate that these things make it less likely that a uke-player would be as equipped to decode songs as compared to a guitar player, and that this leads to less songs being tabulated and uploaded. Pound-for-pound, the numbers would be way off, but taken as a ratio of (real) uploaders to people searching for uploads, the uke ratio must look way different than that of the guitar. Also, for some reason, I say that a guitar player would be less likely to run a simple search rather than going to a trusted source (due to being more of a professional player? serious speculation here based on loose ideas of the comparative profiles of guitar-vs-uke players, I understand).

Mike and Jarvis' reggae-playing Ukulego robots 

Overall, when looking at the cyber-uke-sphere, it just smells like fertile soil for a place like a tab-generating robot to entice hapless players.

And though I may not know music very well, and I don't know how to actually program a robot, it can't be that hard to make a botsite that restates your search via a songtitle-corrector, a lyric-matcher, and a cache of chord names and respective key groupings.

It's just too bad they can't figure out how to actually decode the songs for us instead of just pretending to do it.

One day, Leonard B. Meyers will be proud...
Music, the Arts and Ideas, Leonard B. Meyers, 1967: Music as a Learned Probability System


Sandra -psychosandra- Holmbom

POST-POST SCRIPT
...something else about the chronologically stipulated evolution of the instruments respective to that of the internet...kind of like what happened to the ampersand in English vs. French, but not really.

Technologically-mediated cultural artifacts of both the ampersand and tab-generator software, see below.
The Ampersand
October 2012

Friday, November 2, 2012

Gotcha


How Companies Learn Your Secrets
[aka: how target knows you’re pregnant before you do]
CHARLES DUHIGG, February 16, 2012

modified article:
For companies like Target, the exhaustive rendering of our conscious and unconscious patterns into data sets and algorithms has revolutionized what they know about us and, therefore, how precisely they can sell.

THE BACKGROUND SCIENCE:
Basically, habits help us think less:
An M.I.T. neuroscientist named Ann Graybiel told me that she and her colleagues began exploring habits more than a decade ago by putting their wired rats into a T-shaped maze with chocolate at one end. The maze was structured so that each animal was positioned behind a barrier that opened after a loud click. The first time a rat was placed in the maze, it would usually wander slowly up and down the center aisle after the barrier slid away, sniffing in corners and scratching at walls. It appeared to smell the chocolate but couldn’t figure out how to find it. There was no discernible pattern in the rat’s meanderings and no indication it was working hard to find the treat.

The probes in the rats’ heads, however, told a different story. While each animal wandered through the maze, its brain was working furiously. Every time a rat sniffed the air or scratched a wall, the neurosensors inside the animal’s head exploded with activity. As the scientists repeated the experiment, again and again, the rats eventually stopped sniffing corners and making wrong turns and began to zip through the maze with more and more speed. And within their brains, something unexpected occurred: as each rat learned how to complete the maze more quickly, its mental activity decreased. As the path became more and more automatic — as it became a habit — the rats started thinking less and less.

This process, in which the brain converts a sequence of actions into an automatic routine, is called “chunking.” There are dozens, if not hundreds, of behavioral chunks we rely on every day. Some are simple: you automatically put toothpaste on your toothbrush before sticking it in your mouth. Some, like making the kids’ lunch, are a little more complex. Still others are so complicated that it’s remarkable to realize that a habit could have emerged at all.

Take backing your car out of the driveway. When you first learned to drive, that act required a major dose of concentration, […] Now, you perform that series of actions every time you pull into the street without thinking very much. Your brain has chunked large parts of it. Left to its own devices, the brain will try to make almost any repeated behavior into a habit, because habits allow our minds to conserve effort.

To understand this a little more clearly, consider again the chocolate-seeking rats. What Graybiel and her colleagues found was that, as the ability to navigate the maze became habitual, there were two spikes in the rats’ brain activityonce at the beginning of the maze, when the rat heard the click right before the barrier slid away, and once at the end, when the rat found the chocolate. Those spikes show when the rats’ brains were fully engaged, and the dip in neural activity between the spikes showed when the habit took over. From behind the partition, the rat wasn’t sure what waited on the other side, until it heard the click, which it had come to associate with the maze. Once it heard that sound, it knew to use the “maze habit,” and its brain activity decreased. Then at the end of the routine, when the reward appeared, the brain shook itself awake again and the chocolate signaled to the rat that this particular habit was worth remembering, and the neurological pathway was carved that much deeper.

The process within our brains that creates habits is a three-step loop. First, there is a cue, a trigger that tells your brain to go into automatic mode and which habit to use. Then there is the routine, which can be physical or mental or emotional. Finally, there is a reward, which helps your brain figure out if this particular loop is worth remembering for the future. Over time, this loop — cue, routine, reward; cue, routine, reward —becomes more and more automatic. The cue and reward become neurologically intertwined until a sense of craving emerges. What’s unique about cues and rewards, however, is how subtle they can be. Neurological studies like the ones in Graybiel’s lab have revealed that some cues span just milliseconds. And rewards can range from the obvious (like the sugar rush that a morning doughnut habit provides) to the infinitesimal (like the barely noticeable — but measurable —sense of relief the brain experiences after successfully navigating the driveway). Most cues and rewards, in fact, happen so quickly and are so slight that we are hardly aware of them at all. But our neural systems notice and use them to build automatic behaviors.

Our relationship to e-mail operates on the same [cue-routine-reward] principle. When a computer chimes or a smartphone vibrates with a new message, the brain starts anticipating the neurological “pleasure” (even if we don’t recognize it as such) that clicking on the e-mail and reading it provides. That expectation, if unsatisfied, can build until you find yourself moved to distraction by the thought of an e-mail sitting there unread — even if you know, rationally, it’s most likely not important. On the other hand, once you remove the cue by disabling the buzzing of your phone or the chiming of your computer, the craving is never triggered, and you’ll find, over time, that you’re able to work productively for long stretches without checking your in-box.

PERSONAL ANALYTICS:
Find the customers who have children and send them catalogs that feature toys before Christmas. Look for shoppers who habitually purchase swimsuits in April and send them coupons for sunscreen in July and diet books in December.

In the 1980s, a team of researchers led by a U.C.L.A. professor named Alan Andreasen undertook a study of peoples’ most mundane purchases, like soap, toothpaste, trash bags and toilet paper. They learned that most shoppers paid almost no attention to how they bought these products, that the purchases occurred habitually, without any complex decision-making. Which meant it was hard for marketers, despite their displays and coupons and product promotions, to persuade shoppers to change.

But when some customers were going through a major life event, like graduating from college or getting a new job or moving to a new town, their shopping habits became flexible in ways that were both predictable and potential gold mines for retailers. The study found that when someone marries, he or she is more likely to start buying a new type of coffee. When a couple move into a new house, they’re more apt to purchase a different kind of cereal. When they divorce, there’s an increased chance they’ll start buying different brands of beer.

Consumers going through major life events often don’t notice, or care, that their shopping habits have shifted, but retailers notice, and they care quite a bit. At those unique moments, Andreasen wrote, customers are “vulnerable to intervention by marketers.” In other words, a precisely timed advertisement, sent to a recent divorcee or new homebuyer, can change someone’s shopping patterns for years.

And among life events, none are more important than the arrival of a baby. At that moment, new parents’ habits are more flexible than at almost any other time in their adult lives. If companies can identify pregnant shoppers, they can earn millions.

…able to identify about 25 products that, when analyzed together, allowed him to assign each shopper a “pregnancy prediction” score. (like unscented lotion)

One Target employee I spoke to provided a hypothetical example. Take a fictional Target shopper named Jenny Ward, who is 23, lives in Atlanta and in March bought cocoa-butter lotion, a purse large enough to double as a diaper bag, zinc and magnesium supplements and a bright blue rug. There’s, say, an 87 percent chance that she’s pregnant and that her delivery date is sometime in late August. What’s more, because of the data attached to her Guest ID number, Target knows how to trigger Jenny’s habits. They know that if she receives a coupon via e-mail, it will most likely cue her to buy online. They know that if she receives an ad in the mail on Friday, she frequently uses it on a weekend trip to the store. And they know that if they reward her with a printed receipt that entitles her to a free cup of Starbucks coffee, she’ll use it when she comes back again.

About a year after Pole created his pregnancy-prediction model, a man walked into a Target outside Minneapolis and demanded to see the manager. He was clutching coupons that had been sent to his daughter, and he was angry, according to an employee who participated in the conversation.

Using data to predict a woman’s pregnancy, Target realized soon after Pole perfected his model, could be a public-relations disaster. So the question became: how could they get their advertisements into expectant mothers’ hands without making it appear they were spying on them? How do you take advantage of someone’s habits without letting them know you’re studying their lives?

“We have the capacity to send every customer an ad booklet, specifically designed for them, that says, ‘Here’s everything you bought last week and a coupon for it,’ ” one Target executive told me. “We do that for grocery products all the time.” But for pregnant women, Target’s goal was selling them baby items they didn’t even know they needed yet.

“With the pregnancy products, though, we learned that some women react badly,” the executive said. “Then we started mixing in all these ads for things we knew pregnant women would never buy, so the baby ads looked random. We’d put an ad for a lawn mower next to diapers. We’d put a coupon for wineglasses next to infant clothes. That way, it looked like all the products were chosen by chance.

“And as long as we don't spooky her, it works."

As Pole told me the last time we spoke: “Just wait. We’ll be sending you coupons for things you want before you even know you want them.”

Monday, October 29, 2012

Hidden Economies and the Shifting of Value



"The Clothesline Paradox"
Tim O Reilly [10.4.2012]
Edge Conversations, www.edge.org

Stuart Brand's Clothesline Paradox:
You put your clothes in your dryer, and the energy you use gets measured and counted; you put your clothes on the clothesline, and it disappears from the economy.

On the internet, the value is created somewhere, and captured somewhere else. (The sun creates the value, and your wet clothes capture the value). Tim Berners Lee created value in the internet, but did not capture it. Goldman Sach's did not create value, they captured it.

Free content on the web - Users getting something for nothing:
Actually, most people pay comcast $80 a month for content on the web. And what's more, it's actually comcast who gets the free ride now. They don't have to pay television networks for their content; on the internet, the users create that content. Comcast is getting the free ride, not the users.

Finding meaning in the data is the new value generation.


"Thinking in Network Terms"
Albert-lászló Barabási [9.24.2012]
Edge Conversations, www.edge.org

Call it what you will: Network Science, Human Dynamics, Computational Social Science, Big Data; The question now is not how you collect the data, but how you make sense of it.

Barabasi continues to talk about how understanding networks allows us to get more out of the data.



Further Links:
Tim O Reilly is the founder and CEO of O'Reilly Media, Inc., one of the leading computer book publishers in the world.

O'Reilly Radar: Insights, analysis and research about emerging technologies

ALBERT-LÁSZLÓ BARABÁSI is a Distinguished University Professor at Northeastern University, where he directs the Center for Complex Network Research, and holds appointments in the Departments of Physics, Computer Science and Biology, as well as in the Department of Medicine, Harvard Medical School and Brigham and Women Hospital, and is a member of the Center for Cancer Systems Biology at Dana Farber Cancer Institute.
Barabasi Labs, Center for Complex Network Research, Northeastern University, Boston