Showing posts with label big data. Show all posts
Showing posts with label big data. Show all posts

Thursday, April 10, 2025

Data Attack - It's Personal


Perennial public service announcement that all your data are belong to us:

ParkMobile app update: Deadline for $32.8M data breach settlement is here
Mar 2025, nj.com

As many as 21 million people could be eligible for the settlement, which stems from a data breach that happened in 2021.

ParkMobile, popular app used by many Jersey Shore beach towns that collect parking fees, agreed to the $32.8 million settlement to resolve claims “relating to an unknown actor’s unauthorized access to the Personal Information of ParkMobile App users,” the settlement website said.

Mostly unrelated image credit: AI Art - Sandwich Man 2 - 2024


Ex-cop admits hacking into social media accounts of nearly 20 women, distributing naked pics, officials say
Mar 2025, nj.com

A former Mount Laurel police officer admitted this week to hacking into multiple women’s social media accounts and distributing their nude photos, authorities said.

The investigation began in September 2022. He was a rookie officer with the Mount Laurel Police Department at the time, was arrested on Oct. 21, 2022 and charged with three counts of computer crime, invasion of privacy and two counts of endangering the welfare of a child, investigators said.

As the investigation continued, 18 more women were found to be victimized by him, police said. Investigators determined that all the victims had a student email account through Rowan College in Burlington County, the office said.

Detectives learned that he illegally accessed approximately 5,000 email accounts associated with the college, authorities said. He hacked the accounts from his own personal devices while on duty as a patrol officer, according to the release.

Tuesday, May 9, 2023

The Big Hard


Drinking alcohol brings no health benefits, study finds
Apr 2023, phys.org

You've been told that one drink a day is actually good for your health -- but it's so hard to believe, right? It's good for your heart! It can't be true right?

No, it's not true. We screwed up, for decades, using one bad study after another to support a crazy idea. Can it be true that not one person stopped and said, wait, that sounds too good, let me double check that study. Not until now. 

It's been named "former-drinker bias", and it will be in every public health textbook for the rest of time starting now, as an example of what can go wrong with biostatistics and epidemiological research. 

We heard rumblings of this a while back...

And now this article sums it up pretty well:
  • Former drinkers aren't lifetime abstainers -- For example, many studies tend to place former drinkers in the same group as lifetime abstainers, referring to them all as "non-drinkers," Stockwell said.
  • But former drinkers typically have given up or cut down on alcohol because of health problems, Stockwell said. The new analysis found that former drinkers actually have a 22% higher risk of death compared to abstainers.
  • Their presence in the "non-drinker" group biases the results, creating the illusion that light daily drinking is healthy, Stockwell said.
  • It's called "former-drinker bias"; and the reason it's been hiding in our public health research for decades? 
  • "This is an overview of a lot of really bad studies," Stockwell said. "There's a lot of confounding and bias in these studies, and our analysis illustrates that."

via Canadian Institute for Substance Use Research at the University of Victoria in British Columbia:  Jinhui Zhao et al, Association Between Daily Alcohol Intake and Risk of All-Cause Mortality, JAMA Network Open (2023). DOI: 10.1001/jamanetworkopen.2023.6185



Post Script:
Continuum of Risk
  • 2 standard drinks or less a week -- You are likely to avoid alcohol-related consequences for yourself or others at this level.
  • 3 to 6 standard drinks a week -- Your risk of developing several types of cancer, including breast and colon cancer, increases at this level.
  • 7 standard drinks or more a week -- Your risk of heart disease or stroke increases significantly at this level.

Bonus:
Partially unrelated, but still a good example of why science is hard:
HUGO (Human Genome Organisation) Gene Nomenclature Committee (HGNC), the body that names genes, has changed 27 genes to avoid being confused by Excel's default naming protocols.

For example, SEPT2 is the short name of a gene called Septin 2....
-Scientists rename human genes to stop Microsoft Excel from misreading them as dates
Aug 2020, The Verge

Wednesday, January 11, 2023

Leave No Trace


Sensors can tap into mobile vibrations to eavesdrop remotely, researchers find
Oct 2022, phys.org

They could detect the vibrations of a cell phone's earpiece and decipher what the person on the other side of the call was saying with up to 83% accuracy using an off-the-shelf automotive radar sensor and a novel processing approach (called an "eavesdropping attack").

via Penn State: Suryoday Basak et al, mmSpy: Spying Phone Calls using mmWave Radars, 2022 IEEE Symposium on Security and Privacy (SP) (2022). DOI: 10.1109/SP46214.2022.9833568

Image credit: Thermal Tent - Infrared Imaging Services


AI-driven 'thermal attack' system reveals computer and smartphone passwords in seconds
Oct 2022, phys.org

After users type their passcode, a thermal camera can take a picture that reveals the heat signature of where their fingers have touched the device; the brighter an area appears in the thermal image, the more recently it was touched.

via University of Glasgow: Norah Alotaibi et al, ThermoSecure: Investigating the effectiveness of AI-driven thermal attacks on commonly used computer keyboards, ACM Transactions on Privacy and Security (2022). DOI: 10.1145/3563693


Wednesday, September 7, 2022

It Knows


AKA The Great Recognizer

Neural network can read tree heights from satellite images
Apr 2022, phys.org

We need to be reminded of how powerful data can be when it's f**king massive. It doesn't even have to be related. Like for example I can tell which zip code you grew up in, even the street and maybe even the house, based on really really fine-grained data about your teethbrushing habits, delivered by your smart toothbrush of course.

It sounds crazy, but given enough data, nothing is crazy. 

We don't even have to know how to do it, just give a neural network enough data, and it will figure out the problem for you:

"Since we don't know which patterns the computer needs to look out for to estimate height, we let it learn the best image filters itself."

All you need are some training data, so in this case that means a bunch of trees for which we do know the height. Then we take the (otherwise flat) satellite data, mash it with the height data for the known trees to teach the network, and then unleash the network on the unknowns.  

And why do we want to know how tall trees are? "Because whenever we cut down trees, we release carbon into the atmosphere, and we don't know how much carbon we are releasing."

via ETH Zurich: Nico Lang, Walter Jetz, Konrad Schindler, Jan Dirk Wegner, A high-resolution canopy height model of the Earth. arXiv:2204.08322v1 [cs.CV], arxiv.org/abs/2204.08322

Image source: The 4 Trends That Prevail on the Gartner Hype Cycle for AI, 2021, Gartner, Sep 2021
via: Intelligent Sensing: Enabling the Next “Automation Age” by Marco Cassis of STMicroelectronics at International Solid-State Circuits Conference 2022


Anxious individuals identified by analyzing their walking gait
May 2022, phys.org

The best method for identifying anxious individuals was walking. The team successfully identified people who were anxious with 75% accuracy.

They had to complete a balance test and a two-minute walk while wearing sensors. Based on these data, the team determined the young people who report being anxious walk in a way that's very similar to older adults who are fearful of falling. They find that young, anxious adults are constantly scanning for threats from side to side while walking and have trouble turning. The researchers also reported that anxious people have worse balance than those who are anxious.

via Clarkson University: Maggie Stark et al, Identifying Individuals Who Currently Report Feelings of Anxiety Using Walking Gait and Quiet Balance: An Exploratory Study Using Machine Learning, Sensors (2022). DOI: 10.3390/s22093163


Using electric signals from human brains, new software can perform computerized image editing
Jun 2022, phys.org

"All the existing software has been previously trained with labeled input. So, if you want an app which can make people look older, you feed it thousands of portraits and tell the computer which ones are young, and which are old.

Here, the brain activity of the subjects was the only input.

This is an entirely new paradigm in artificial intelligence—using the human brain directly as the source of input."

"All the existing software has been previously trained with labeled input. So, if you want an app which can make people look older, you feed it thousands of portraits and tell the computer which ones are young, and which are old. Here, the brain activity of the subjects was the only input. This is an entirely new paradigm in artificial intelligence—using the human brain directly as the source of input."

But alas:
"Collecting individual brain signals does involve ethical issues..."

We are the data, and we have lots of it.

Be a real shame if someone were to save our brainwaves and then sell them on the data market, where they could in turn unintentionally feed back to us our own biases after having been amplified by some behaviorally-exploitative algorithm...

via University of Copenhagen and University of Helsinki: Brain-Supervised Image Editing. Keith M. Davis III et al. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun 2022. 


People with similar faces likely have similar DNA
Aug 2022, phys.org

They recruited human doubles from the photographic work of François Brunelle, a Canadian artist who has been obtaining worldwide pictures of look-alikes since 1999. They obtained headshot pictures of 32 look-alike couples. 

Physical traits such as weight and height, as well as behavioral traits such as smoking and education, were correlated in look-alike pairs. Taken together, the results suggest that shared genetic variation not only relates to similar physical appearance, but may also influence common habits and behavior.

Eventually the all-seeing omnibot will be able to create a complete living human population gene network based only on our faces (and it will predict our habits and behaviors too).

via Barcelona Supercomputing Center and Josep Carreras Leukaemia Research Institute Barcelona: Manel Esteller, Look-alike humans identified by facial recognition algorithms show genetic similarities, Cell Reports (2022). DOI: 10.1016/j.celrep.2022.111257


Algorithm predicts crime a week in advance, but reveals bias in police response
Jul 2022, phys.org

Algorithm that forecasts crime by learning patterns in time and geographic locations from public data on violent and property crimes. The model can predict future crimes one week in advance with about 90% accuracy.

Not like that!
In a separate model, the research team also studied the police response to crime by analyzing the number of arrests following incidents and comparing those rates among neighborhoods with different socioeconomic status. They saw that crime in wealthier areas resulted in more arrests, while arrests in disadvantaged neighborhoods dropped. Crime in poor neighborhoods didn't lead to more arrests, however, suggesting bias in police response and enforcement.

via University of Chicago: Ishanu Chattopadhyay, Event-level prediction of urban crime reveals a signature of enforcement bias in US cities, Nature Human Behaviour (2022). DOI: 10.1038/s41562-022-01372-0


AI reveals unsuspected math underlying search for exoplanets
May 2022, phys.org

I forget about the singularity sometimes, but it's happening right now; we're already in the singularity:
Artificial intelligence (AI) algorithms trained on real astronomical observations now outperform astronomers in sifting through massive amounts of data to find new exploding stars, identify new types of galaxies and detect the mergers of massive stars, accelerating the rate of new discovery in the world's oldest science.

But that's already been the case, now it appears the AI has discovered "unsuspected connections hidden in the complex mathematics arising from general relativity". 
The algorithm figured out new rules of gravitational microlensing:
"I argue that they constitute one of the first — if not the first — time that AI has been used to directly yield new theoretical insight in math and astronomy."-Joshua Bloom, UC Berkeley professor of astronomy

"Keming's machine learning algorithm uncovered this degeneracy that had been missed by experts in the field toiling with data for decades. This is suggestive of how research is going to go in the future when it is aided by machine learning, which is really exciting." -Scott Gaudi, professor of astronomy at Ohio State

via University of California - Berkeley: Keming Zhang et al, A ubiquitous unifying degeneracy in two-body microlensing systems, Nature Astronomy (2022). DOI: 10.1038/s41550-022-01671-6


Post Script:
"AI software has collaborated with mathematicians to successfully develop a theorem about the structure of knots, but the suggestions given by the code were so unintuitive that they were initially dismissed. Only later were they discovered to offer invaluable insight. The work suggests AI may reveal new areas of mathematics where large data sets make problems too complex to be comprehended by humans."
-DeepMind AI collaborates with humans on two mathematical breakthroughs, Matthew Sparkes, New Scientist, Dec 2021 [link]

They're talking about knots. Instead of feeding it facebook photos or product reviews, they gave it math problems...about knots. The "connection between algebraic and geometric invariants of knots".

And why knots? Because theories about knots can also be applied to quantum field theory. Don't ask me why, but they do, and it's called topology.

via Google DeepMind: Advancing mathematics by guiding human intuition with AI. Davies, A., Veličković, P., Buesing, L. et al. Nature 600, 70–74 (2021). DOI: 10.1038/s41586-021-04086-x


Post Post Script, On Synthetic Data:
'Fake' data helps robots learn the ropes faster
Jul 2022, phys.org

How to be a human:
For the rope-looping simulation and experiment, Mitrano and Berenson expanded the data set by extrapolating the position of the rope to other locations in a virtual version of a physical space—so long as the rope would behave the same way as it had in the initial instance. Using only the initial training data, the simulated robot hooked the rope around the engine block 48% of the time. After training on the augmented data set, the robot succeeded 70% of the time.
via  University of Michigan: Data Augmentation for Manipulation, arXiv:2205.02886v3 [cs.RO]  https://doi.org/10.48550/arXiv.2205.02886


Thursday, August 11, 2022

Greatest Retronym in History


When it comes to AI, can we ditch the datasets?
Mar 2022, phys.org

Synthetic fucking data. They're making synthetic data to train the robots. And that makes us analog data. Me and you, our faces, our fingerprints, our gaits, gestures, voices (and most especially our consumer behaviors), are analog, starting now.

First there was the acoustic guitar, then dairy milk ffs. Hopefully, when we finally cede control to the omnibot envelope, we don't go the way of the flip-phone. 

via MIT: Paper: Generative models as a data source for multiview representation learning. openreview.net/pdf?id=qhAeZjs7dCL


Physiological signals could be the key to 'emotionally intelligent' AI, scientists say
Apr 2022, phys.org

You got any more of that analog data?
They're coming for your sweat, your biodata. You are the training set for the artificial humans of the future. 

via Japan Advanced Institute of Science and Technology JAIST: Shun Katada et al, Effects of Physiological Signals in Different Types of Multimodal Sentiment Estimation, IEEE Transactions on Affective Computing (2022). DOI: 10.1109/TAFFC.2022.3155604

Image credit: Jared Michael

Friday, March 25, 2022

Ubiquitous Intelligence


Solving the 'big problems' via algorithms enhanced by 2D materials
Jan 2022, phys.org

Imagining the future is hard; you almost never get it right. You just can't see what's not there.  But in this case, a glimpse reveals itself -- every "computer" will be designed for a specific algorithm. There won't be a "new" way of making a computer, or of making a one-size-fits-all computer that's faster or better. The very idea of a one-size-fits-all computer is what makes it hard for us to see the future. Before the electric guitar was invented, the acoustic guitar didn't exist, it was just called a guitar. (Like smartphones and dumbphones.)

Eventually, you won't have an advanced computer that can run different algorithms better, instead the computer and the algorithm will be one, and therefore there will be as many types of computers as there are algorithms. Like the Cambrian explosion, but different. 

The "combinatorial optimization problem" they're solving here is also referred to as the "traveling salesman problem", or the "design an optimal transit system based on the terrain of the region, distribution of the population, existing routes, etc.", or the "use a living slime mold computer to design an optimal transit system" problem. 

The reason it's so hard for our current algorithms to do the optimization problem is because computers as we know them today are still based on a design from the 1940's. The problem isn't so much because they're old, it's because we don't work with data the same way we used to. We have a lot more data than we used to, and integrating it all at the same time is hard for today's computers.

I like to think of it simply as a problem where your database has as many columns as it does rows. This is what happens when you try to categorize smells based on the names we call them. On one axis you have all the smellable molecules there are (veritably infinite), and on the other, you have all the attributes you can give to any one of the molecules (physical dimensions, descriptions, names, autobiographical physiodata that your body associates with the molecule, which is also veritably infinite). You would then have a database, a spreadsheet of infinite cells. It's hard to work with something that big. 

But we don't have to do things like that anymore. Instead, we can use slime mold, or we can design "new computers" that combine information storage and computing into the same thing. This sounds a lot like a neuromorphic computer, by the way.

via Pennsylvania State University: Amritanand Sebastian et al, An Annealing Accelerator for Ising Spin Systems Based on In‐Memory Complementary 2D FETs, Advanced Materials (2021). DOI: 10.1002/adma.202107076

Image credit: Flows of individuals across the Greater Boston area, Guangyu Du at Sante Fe Inst, 2021

Post Script:
And how they do it? A form of "in-memory computing" based on simulated annealing, where atoms reorganize themselves and then crystallize in the lowest energy state. Sounds a lot like 2-D metamaterials, BECs and quantum crystallography. Putting it all together. 

Notes:
Using a 'virtual slime mold' to design a subway network less prone to disruption
Feb 2022, phys.org

A model, no slime needed.

via University of Toronto: Raphael Kay et al, Stepwise slime mould growth as a template for urban design, Scientific Reports (2022). DOI: 10.1038/s41598-022-05439-w

Monday, March 14, 2022

Look Mom No Data


AKA From Deep Learning to Deep Reasoning

DRNets can solve Sudoku, speed scientific discovery
Sep 2021, phys.org

You can teach a machine to recognize a dog by showing it 1,000 pictures of dogs, Gomes said, but scientific discovery is not like that.

"You are not going to have lots and lots of labeled data," she said. "And in general, the examples you have are not exactly what you are looking for, but then you reason about what you know scientifically about the domain, and you can infer new knowledge."

Key to DRNets is the idea of an "interpretable latent space." Basically, it gives DRNets the ability to reason about the constraints of the domain—in this case materials science—from input data.

They started with Sudoku -- de-mixing overlapping handwritten Sudoku puzzles—grids. The computer had to separate the puzzles into two solved Sudokus, without any training data, which it was able to achieve with close to 100% accuracy.

The researchers then put DRNets to work on a real-world problem: automating crystal-structure phase mapping of solar-fuels materials, using X-ray diffraction (XRD) patterns. Crystal-structure phase mapping involves separating the source XRD signals of the desired crystal structures from "noisy" mixtures of XRD patterns, a task for which labeled training data are typically not available. ... DRNets was able to identify and separate a total of 13 crystal phases (single-phase materials) in 19 unique mixtures of the single-phase materials. ... DRNets' findings, verified using manual analysis, enable the discovery of complex mixtures of crystalline materials that convert solar energy into storable solar chemical fuels.

via Cornell University: Di Chen et al, Automating crystal-structure phase mapping by combining deep learning with constraint reasoning, Nature Machine Intelligence (2021). DOI: 10.1038/s42256-021-00384-1


Wednesday, October 2, 2019

On the Multi-Dimensionality of Cultural Communication


Facebook 'labels' posts by hand, posing privacy questions
May 2019, Reuters

Facebook uses only five dimensions to categorize your pictures. Of these, we have these: 1. Subject (food, person, animal), 2. Occasion (day at the office, 1st birthday party), 3. Intention (plan, inspire, joke). The labeling is done by hand, in order to train machines. Meat-handlers they're called in the industry. Just kidding I made that up; but that's what they are -- because our machines are too stupid right now to be able to do this.

The problem is that humans are also too stupid. Rephrase that -- it's not that humans are stupid, it's the wrong humans being used. The Big F-Book uses meatmen in Romania and the Philippines. Now I'm not sure if I'm getting this right, but it sounds like some human flesh engines from one culture are interpreting the actions of another culture a half a world away.

The problem is when things get lost in translation. If we don't get the cultural nuances right, the resulting data will be messed up. Imagine there is some little quirk, a little difference or misunderstanding between the Filipino labeler and the American poster that labels all x-posts as y. And that error gets scaled up until a huge mistake is made when screening your background for some dystopian automated system that you really want to be a part of, or not, like the criminal justice system for example.

We can't ignore these differences. They may seem small, but they get scaled by the millions. At 150mph, the tiniest pebble will throw your motorcycle right off the road.

Post Script:
Talking about multi-dimensionality, try categorizing smells.

Friday, July 12, 2019

Fuhgeddaboutit


The Right to Be Forgotten was an interesting turn. Something tells me that Forgetting will be big business pretty soon. As we reach a point of data saturation, we're now trying to come up with ways to not have "absolute data".

Amazon digital assistant Alexa gets new skill: amnesia
May 2019, phys.org

Microsoft deletes massive face recognition database
Jun 2019, BBC News

Interesting idea - deleting an online database:

"You can't make a data set disappear," Adam Harvey from the Megapixels site told Engadget."Once you post it, and people download it, it exists on hard drives all over the world."


Post Script
Researchers erase fearful memories in mice
AAAS, Aug 2014


Sunday, July 1, 2018

Sans Agency Humans

aka Sociothermodynamics


People get all bent out of shape thinking about the eminent takeover of artificial intelligence. Personally I think we're already robots, or rather, we've always been.

Bees in a hive, wolves in a pack, gas molecules in a prescribed volume. Do we really make decisions or does something else do it for us? And I don't mean God, I mean physics.

A nice string of headlines surfaced lately that alludes to this idea of humans being driven by forces well beyond our control:

Research finds tipping point for large-scale social change
Jun 2018, phys.org

Roughly 25% of people need to take a stand before large-scale social change occurs. This idea of a social tipping point applies to standards in the workplace and any type of movement or initiative.

How physics explains the evolution of social organization
Jun 2018, phys.org

A scientist at Duke University says the natural evolution of social organizations into larger and more complex communities that exhibit distinct hierarchies can be predicted from the same law of physics that gives rise to tree branches and river deltas. 
[Author] outlines how these seemingly disparate phenomena are actually connected through the constructal law of evolution in nature. Penned by Bejan in 1996, the law states that for a system to survive, it must evolve over time to increase its access to flow. [...] and the true nature of an innovation is simply a local design change that increases the efficiency of the distribution of a resource to the entire population. While the individual innovator may benefit greatly from the idea, the entire community also gains better access to that resource, which serves to reduce overall inequality.

Study of Google search histories reveals relationship between anti-Muslim and pro-ISIS sentiment in U.S.
Jun 2018, phys.org

Researchers suggest, targeting [Islamic] groups in countries such as the U.S. might be causing home-grown radicalization to occur. 
[They] found a common theme—in low income communities where there were a lot of anti-Muslim searches, there were also a lot of searches by people looking for more information about radical Islamic groups. Such communities, the researchers further noted, tended to be homogeneous in nature, mostly white, with few people of color. People from the Middle East, they point out, stand out in such communities. This finding, they claim, suggests that anti-Muslim activities such as discrimination and being targeted by government officials might actually be pushing some of those targeted people toward becoming extremists.

Adding a bit of moderation, we should note that humans lack agency in this context the way that smoking one cigarette takes 0.05 seconds off your life. Assertions like this come from absolutely huge, unimaginable numbers of people and variables working in concert and across time.

Our little brains can't handle it. This is why it's so hard to connect that one cigarette to lung cancer (especially in the face of nicotine addiction) and also why many if not most of us can't handle the possibility that neither we nor an omnipotent entity have control. It is instead the interaction of myriad forces, generating every-increasingly complex arrangements. Coincidentally, it is Liebniz bday (thanks google doodle); he is the guy that said things don't exist in themselves but only in their relation to others - an idea that has yet to be fully incorporated into our universal model of this thing we live in.

image source: valve steam controller

Sunday, June 3, 2018

Vision, Accuracy, and the Right Brain


In a TED talk I can't seem to forget, Iain McGilchrist, in RSA fashion, animates the two sides of ourselves. These are the two sides of our brains, the two hemispheres. I don't think I need to explain the Left Brain - Right Brain distinction on account of its general popularity. I will only say that this is not an absolute thing; it's a heuristic for understanding how our complicated heads run all that bugged out code up there.

The part of McGilchrist's talk I can't forget lies in his premonition that humans have been heading in the Left direction since as far back as we can remember, but it may soon be time for the pendulum to turn the other direction. It's hard to believe. We can't have megacities without the Left Brain. No spaceships, no biodomes, no nanofabricated body modifications.

Or can we? If you read about contemporary advances in artificial intelligence, you might be thinking that it's quite possible.

Much of the new things happening in this field (which are actually old ideas running on new hardware) are based on a pretty revolutionary paradigm.  I'm referring here to deep learning neural nets and more generally the idea of machine learning. This semantically refers to an approach to computing that uses a kind of brute force instead of a superintelligent program. It's kind of like a Wisdom of the Crowds thing plus computers; instead of one thousand guesses, there's 100 trillion guesses. (I say 'a kind of brute force' because in other instances, some would could the traditional method a brute force of computation.)

So the answer they come up with isn't exact, it's approximate. But when you have to query petabytes of data, you can no longer expect an exact answer.

Check out for example this new thing where scientists Frankenstein a neural net and a cell phone together to makeshift a microscope as powerful as one in a high grade laboratory. We no longer need to see every micropixel to get a clear picture of what we're looking at. Instead we teach an algorithm to see, and it does something that in theory is similar to what our brain does when we see. We make stuff up, we fill in the dots, we approximate. (In actuality, the algorithm is taught not how to see, but how to learn to see.)

We live in this post-truth world, right? Facts don't matter; belief matters. Being right is not as important as convincing people that you're right. That's a tangent. But we do live in a world of Big Data, so big we really don't know what to do with it all. We simply can't make computers powerful enough to handle all of it. But we have this new approach that can scale-up. It approximates, which is not what we're used to, but it does reach the scale of Big Data.

Have we reached this inflection point where our technology is starting to act more like our brain (the whole brain, not just the left side)? Is there where we find out, after hundreds of years (after the Enlightenment / Scientific Revolution) that absolute certainty is not the ultimate goal in all knowledge-gathering endeavors? Just speculating here, but it sure seems like we're headed for a future that looks more like a wet biological mess than a crystallized spreadsheet.

Notes

Deep learning transforms smartphone microscopes into laboratory-grade devices
May 2018, phys.org

Iain McGilchrist, The Divided Brain, 2011
this is a TED talk based on his book

Bicameralism 
Julian Jaynes, The Origin of Consciousness and the Breakdown of the Bicameral Mind, 1976

Tuesday, August 15, 2017

Eyes on the Street


Computer 'anthropologists' study global fashion
Aug 2017, phys.org

What is the world wearing?

These scientists are using a deep learning object recognition program to discover visual patterns in clothing and fashion across millions of images of people worldwide and over a period of many years. They detected attributes like color, sleeve length, presence of glasses or hats, etc. (They end up filtering for only waist up photos). They ask questions such as, "How is the frequency of scarf use in the US changing over time?" or "For a given city, such as Los Angeles, what styles are most characteristic of that city."

The objective of this research is ultimately to "provide a look into cultural, social and economic factors that shape societies and provides insights into civilization."

Dashed lines mark Labor Day. Who said Americans don't like conformity?

via Cornell University: StreetStyle: Exploring world-wide clothing styles from millions of photos. arXiv. arxiv.org/abs/1706.01869

I imagined that stuff like this is already happening all over the place, in all kinds of other fields, and being integrated into global policy decisions and bottom-line business calls alike. But, this is not the case; this is still just the beginning. One thing I caught from this, some digital era common sense - Google Trends results for "scarves" peak right before they do on Instagram, because, presumably, people are searching for the thing, then they buy it, then they take pictures of themselves wearing it.

Post Script
These are the real people, not the algorithms, that analyze and predict the world of fashion:
Color Conspirators, Network Address

Friday, November 25, 2016

Still Awaiting Omniscience


global-monitoring systems:

"In new research published Thursday in the journal Science, Northeastern network scientist David Lazer and his colleagues analyzed the effectiveness of four global-scale databases and found they are falling short when tested for reliability and validity.

The fully automated systems studied were the International Crisis Early Warning System, or ICEWS, maintained by Lockheed Martin, and Global Data on Events Language and Tone, or GDELT, developed and run out of Georgetown University. The others were the hand-coded Gold Standard Report, or GSR, generated by the nonprofit MITRE Corp., and the Social, Political, and Economic Event Database, or SPEED, at the University of Illinois, which uses both human and automated coding.

"It's so easy for us as humans to read something and know what it means," says Lazer. "That's not so for a set of computational rules."

The authors suggest that reliable data-tracking systems can be used to build models that anticipate the escalation of conflicts, forecast the progression of epidemics, or trace the effect of global warming on the ecosystem."

Using Big Data to monitor societal events shows promise, but the coding tech needs work
phys.org, Oct 2016
http://phys.org/news/2016-09-big-societal-events-coding-tech.html


image source

Thursday, September 26, 2013

The Predictionary



See below: the methodology alone is a complete mindfuck, and should provide some sense of the nature of scientific studies in the age of big data.

Get ready...

Researchers say readers' identities can reveal much about content of articles
Aug 12, 2013
modified article:

Articles that people share on social networks can reveal a lot about those readers, but a  new study reverses the proposition: What can be learned about an article from the attributes of its readers?

To find out, the CMU researchers, along with colleagues at the University of Washington, analyzed almost 3 million news articles and the public profiles of the people who shared those articles on Twitter.

This enabled them to generate a few thousand "badges" that characterized the content of the shared news articles and also could be used to analyze any subsequent article, including those that had never been shared or even read.

In order to train their model, the team began by looking at three months of tweets—from September of 2010, 2011 and 2012—and selecting those that included links to mainstream news articles and came from a user who had filled out a Twitter profile.

[collect major news outlet's articles that have been tweeted about] 

Each news article was then downloaded and the most meaningful, unique words were extracted, creating a "bag of words" for each article; similar to a visual word cloud, these bags give greater weight to more important words. Likewise, from each user's Twitter profile, a set of descriptive words, or badges, was extracted.

[each article, as well as each twitter-user-profile, gets a weighted wordcloud]

By comparing the bags of words with badges from the people who shared the articles, the researchers were able to create a dictionary that associated each badge with its characteristic words. For example, people who self-identify with the music badge in their profiles are likely to share articles with words such as "band," "album" and "song." Different dictionaries were created for each year to compensate for interests or topics that change over time. These dictionaries were then used to encode new articles, leading to a document representation based on attributes of potential readers.

[the article wordclouds and the user-profile wordclouds of the users who tweeted those articles are cross-correlated to create a "dictionary", or rather a "predictionary", if you will, that predicts who will share what, or what will be shared by whom]

Case Study Example:
New York Times columnist Maureen Dowd had readers who tended to be progressive. This association was notable because Dowd never explicitly uses the word "progressive" in the articles analyzed by the researchers. Rather, the algorithm detected that the words Dowd uses in these articles correspond to the type of content self-described progressives tend to share on Twitter.
-Carnegie Mellon University