Showing posts with label probability. Show all posts
Showing posts with label probability. Show all posts

Friday, April 29, 2022

Baselines, Biases and Big Data Problems - The Secrets of Statistics and the Magic of Metrology



The paradox of big data spoils vaccination surveys
Dec 2021, phys.org

Image credit: Labyrinth by Liqen, Miami 2011

So guess what -- big data is really good at minimizing sample size errors, but also at magnifying systematic biases such as nonresponse bias like how vaccinated people are more likely to respond and marginalized groups less so -- in other words, bad data. Yet because the sample size is so large, we are tricked even more into thinking it must be good data. "Biases in the data get worse with bigger sample size." -phys.org

Here's an example: "Two in 10 respondents did not have a college degree, compared with four in 10 of all U.S. adults—and race and ethnicity—the fraction of black and Asian respondents was only half of what it is in the general population."

"Worse than no survey at all" they say.

Thanks Facebook, but you can keep your invasive mass surveillance system of the entire American population to yourself. 

via Harvard: Seth Flaxman, Unrepresentative big surveys significantly overestimate US vaccine uptake, Nature (2021). DOI: 10.1038/s41586-021-04198-4


On research methods and data quality:

Meng said he began thinking about the problems posed by big data during a visit to Harvard a decade ago by a U.S. Census Bureau official. The official met with a group of statisticians and asked them about the handling of data sets that were becoming available covering large percentages of the U.S. population. Using the hypothetical example of tax data collected by the IRS, he asked whether the statisticians would prefer a sample covering 5 percent of the population that they knew was representative of the larger population or IRS data that they weren't sure was representative but covered 80 percent of the population. The statisticians chose the 5 percent. "What if it was 90 percent?" the Census Bureau official asked. The statisticians still chose the 5 percent, because if they understood the data, their answer would likely be more accurate than even a much larger set with unknown biases.

And now for the hard stuff:
Completely unrelated image; I'm just collecting pictures from science articles of people holding vials in their fingers: Wastewater Filtration, Pacific Northwest National Laboratory and Andrea Starr, 2022

Drinking alcohol to stay healthy? That might not work, says new study
Nov 2021, phys.org

This is such a great example of how statistics works (and how it doesn't). The key word is "baseline".

We've been told for years, forever?, that one glass of red wine a day is not only not-bad, it's actually good for you. In health-speak, they call that "protective".

But this study shows us that there is no group of people that we can use as a baseline, who ---doesn't--- drink and yet who also ---doesn't--- have other health issues caused by having a history of substance abuse. In other words, when the only people who don't drink are people who are recovering from being addicted to alcohol (gross exaggeration), it means you can't find a baseline for a "normal" person. And if you can't find a baseline, then you can't measure anything. 

The majority of the alcohol abstainers at baseline were former alcohol consumers and had risk factors that increased the likelihood of early death. Former alcohol use disorders, risky alcohol drinking, ever having smoked tobacco daily, and fair to poor health were associated with early death among alcohol abstainers. Those without an obvious history of these risk factors had a life expectancy similar to that of low to moderate alcohol consumers. The findings speak against recommendations to drink alcohol for health reasons.

I tell my friends this story, and they ask the first question, a good question -- what about people like Seventh Day Adventists, or Muslims, they don't drink, why can't we use them? And the answer is the reason why health science is hard.

The reason you can't use groups of people who don't drink, is because that group would likely not represent the much larger group of, let's say, all the people in America. It doesn't even matter if it's a small group; making it bigger won't help. It doesn't work because it doesn't match. You can't compare the two groups because the people aren't the same. 

Another way we see this, and one which is becoming more evident to those who can fix it, is how certain groups of people (like undocumented immigrants) are under-represented in the data. If the majority of the datapoints are White, Christian and middle class, and you're none of those, then it's possible that the data is not relevant to you. Your baseline isn't represented, so whatever health effects you're trying to measure, they aren't being compared to someone like you. 

via Public Library of Science: John U, Rumpf H-J, Hanke M, Meyer C (2021) Alcohol abstinence and mortality in a general population sample of adults in Germany: A cohort study. PLoS Med 18(11): e1003819. doi.org/10.1371/journal.pmed.1003819

Partially Related:
This is also a correlate to the story of how lead was discovered in the air -- a geochemist was trying to do such sensitive work (to measure the age of the Earth!) that he kept picking up extra lead in his results, and could not figure out where it was coming from. But it turns out that it was in the air, having been vaporized in the internal combustion engines in our cars. And that story implies that until then, all other experiments being done were "wrong" because they didn't exclude the excess lead from the otherwise "normal" background. 

Last One -- Palmar Sweating:
During a nuclear war scare (1950's), all experiments into palmar sweating at a research institute had to be abandoned because the base level of the response had become so abnormal that the tests would have been meaningless.  (p188)
-The Naked Ape, Desmond Morris, 1967

Monday, April 4, 2022

Future Forecasting


The AI forecaster: Machine learning takes on weather prediction
Jan 2022, phys.org

Standard models are still good for the 2-3 week range, but this new deep learning model is on par with them for the 4-6 week range. If anyone read that scene in Neal Stephenson's Terminal Shock where they were predicting an extreme weather event 3 weeks out with a secret Chinese supercomputer -- this is what he's imagining. 

via American Geophysical Union: Jonathan A. Weyn et al, Sub‐Seasonal Forecasting With a Large Ensemble of Deep‐Learning Weather Prediction Models, Journal of Advances in Modeling Earth Systems (2021). DOI: 10.1029/2021MS002502


Study finds US flood damage risk is underestimated
Feb 2022, phys.org

Interesting example of how probability, statistics, and predictive analytics works -- 

The actual flood damage reports they used to "train" the models were publicly available reports from NOAA made between December 2006 and May of 2020. Compared with recent FEMA maps downloaded in 2020, 84.5% of the damage reports they evaluated were not within the agency's high-risk flood areas. The majority, at 68.3%, were located outside of the high-risk floodplain, while 16.2% were in locations unmapped by FEMA.

When they ran their computer models to determine flood damage risk, they found a high probability of flood damage for more than 1.01 million square miles across the United States, while the mapped area in FEMA's 100-year flood plain is about 221,000 square miles. Researchers said there are factors that could help explain why the differences were so large, including that their machine-learning-based model assessed damage from floods of any frequency, while FEMA only includes flooding that would occur from storms that have a 1% chance of happening in any given year [100 year storms].

-- Now remember, here in New Jersey for example, one of the fastest changing climate regions in the world, we had two 500-year storms in two years, one of them a flooding event, the other wind. I'm pretty sure the floods of September 2021 were a 100- if not 500-year storm. All in 10 years. 

Totally unrelated image credit: Fractal Forums, Christmas Ornament, 2019


Post Script, on Predictive Analytics:
Algorithm can predict possible Alzheimer's with nearly 100 percent accuracy
Sep 2021, phys.org

via Kaunas University of Technology: Modupe Odusami et al, Analysis of Features of Alzheimer's Disease: Detection of Early Stage from Functional Brain Changes in Magnetic Resonance Images Using a Finetuned ResNet18 Network, Diagnostics (2021). DOI: 10.3390/diagnostics11061071

Sunday, December 8, 2019

Artificial Impressionism


Fake news via OpenAI - Eloquently incoherent?
Nov 2019, phys.org

Robots slowly taking over. Give them a sentence and they can now keep it going for a few more sentences, but after that it gets stupid.

So you can give it a fake headline, and it will generate the first line of the story, but after that things will start to fall apart.

Good thing the targets for engineered memetic propagation are not trying to read past the first sentence!

Post Script
Researchers develop a method to identify computer-generated text
July 2019, phys.org



In the above 3 images, the first is a chunk of text written by a robot (most of the words are green, with a few yellows sprinkled in), the second is a real New York Times article (only half is green, the rest is yellow, with some red, and a sprinkle of purple) and the third picture is a clip from "the most unpredictable human text ever written", James Joyce's Finnegan's Wake (the colors green, yellow, red, purple are all evenly distributed about the page).

Green words are very predictably the next word. Yellow words are less likely to show up after the word they show up after. And red and purple are for when the next word is something you absolutely did not expect.

Because text-writing algorithms today use a statistical correlation program based on a compendium of written language (so they know what words typically occur together) the output of such algos will tend to look like the topmost image with all green words. The algos can't think for themselves, they can't 'come up with' new stuff, and they can't be unpredictable. The whole point of writing an algorithm to do this is to prescribe what it's going to do in advance, i.e., it's predictable.

Anyway, soon we won't be writing our robots to write like that. They'll use less predictable programs to generate their text, with unpredictability and random association thrown in there on purpose.

Sunday, January 27, 2019

Okay To Be Redhead


In the largest genetic study of hair colour to date (350,000), a team at Edinburgh University has discovered eight previously-unknown genetic differences between redheads and non-redheads.

The only reason I  mention this is to add a phrase that really strikes me lately -- the redheaded stepchild.

I had been hearing it here and there, but within this past year, it was in the headline of a news article in a local newspaper, I believe it was like 'New Jersey being treated like a red-headed stepchild.'

The fact that such a phrase was being used in a (relatively) reputable news source seemed borderline outrageous to me. I am not necessarily a proponent of political correctness, unless of course you're simply referring to common courtesy. But in today's world I thought it was strange that anyone would use, in a derogatory way, a term that is derived from the way someone looks.

But then I realized two things: 1. The current President of the United States is a redhead (no?), and of the two political affiliations that would be most likely to have a problem with such derogatory phrases, the typically offended would also likely be unified in their dislike for the President, and therefore more willing to give this one a pass. And 2. This term seems to be more related to probability than anything else, because of the simple fact that redheads are rare.


Redheads are rare, and if neither you nor your husband, nor any of your other kids have red hair, well it sure looks like you've been shagging the mailman. Hence the two terms - redhead and stepchild - have a lot in common.

Lastly, it sure seems like this term, in its instantiation at turn-of-the-century America, would typically have been delivered with a curl of the lip (contemptuously), at least by the father of said stepchild, specifically because the Irish were heavily discriminated against AND they tended more to have red hair. AND they tended more to be mailmen.jk

Post Script:
The full phrase that puts this in perspective is "As welcome as a red-headed stepchild." It's not so much that the -child is unwelcome, just what they represent, which is adultery, and with someone of a lower class. And so really it's the mother of the child that is, or was, the target of disrepute. Nowadays, according to the way we typically hear it, the child is the target.

Notes:
Gene study unravels redheads mystery
Dec 2018, BBC



Tuesday, July 10, 2018

Post Script

Max Ernst and the rest of the Surrealists experimented with 'automatic generation' a hundred years ago.

I'm reading an article here about how we're now using 'robot-generated script' to make things funny. Because, you know, robots are stupid, and we like to laugh at stupid things.

You give a script-writing robot a thousand Seinfeld episodes and ask it to make a Seinfeld episode. And when it messes up, we laugh.

I'm saying all this half tongue-in-cheek. Don't get me wrong, a lot of this stuff is funny. Maybe these smarty pants experimenting with neural  nets can give you plenty of examples of what I'm talking about.

It's funny when a computer screws up. It's funny when anyone screws up. I had a classmate in third grade who wore yellow-tinted stonewash jeans, and I remember making fun of him and getting in trouble for it. The stonewash was right on for that time in the world of fashion, but the yellow not so much. Things have to be messed up to be funny, but not too messed up.

There's a good formula for funny (and a good graph too) which says the level of funniness in a joke is a function of the probability of the punchline vs your expectations. Researchers exploring 'creative AI' look at the novelty vs the quality, because to be creative we have to be new, novel, unexpected, but not completely out of the ballpark.

My friend in 3rd grade got the stonewash right, but not the yellow dye. A trained neural net (I call them all robots for short) gets most of the material right, but once in a while it throws in there something crazy (something wrong) and we laugh.

The part where things get tricky is when we stop to consider what  we're laughing at - is it the abstracted novelty of the output, or is it that we've assigned agency to the network and are now making fun of it for messing up.

I'm just saying, we might not want to get into the habit of poking fun at these things - not because they will one day retaliate and destroy us, but because they are a reflection of ourselves.


image source: Max Ernst 1937 L'Ange du Foyer - Engel des Kamins

Was That Script Written By A Human Or An AI? Here’s How To Spot The Difference
Jun 2018, Futurism

Anatomy of a Joke
2012, Network Address

Botnik is a community of writers, artists and developers using machines to create things on and off the internet.


Thursday, March 7, 2013

Seeker to Uploader Ratios and Botsites


Sandra -psychosandra- Holmbom

It's tough sometimes to transpose tabs for the uke, believe it or not the majority of uke tabs for certain tunes are software-generated and often wrong. Something about the different ratios of uke-playing uploaders to uke-learning tab-seekers.

Is it because advertisers are more likely to generate botsites for services with higher seeker-to-uploader ratios?



I play the guitar. I also play the ukulele. I don't have much of an idea what I'm doing when I'm trying to figure out how to play a song that I like, as I've always been able to reference a guitar magazine, or because I've always tried to play pretty easy songs with simple chords.

Nowadays, I like to play the uke, and I like to play 7th chords, 9th chords, etc. When I want to learn a song on my uke, I look for online tabs, as I'm sure everyone does. In so doing, I've found that, for certain kinds of songs - typically older songs, like prior to the rise in popularity of the uke (~2010's) - the tabs that can be sourced-up are way off, if not totally wrong. What I usually end up doing, is to look for guitar tabs, and transpose them (which isn't very easy for someone like myself who never really learned music properly). Luckily, I learned some basic music theory in my initial years of guitar playing, and it's been relatively easy to augment that learning, at least enough to meet my needs of detecting the accuracy of my own transpositions.

*The ukulele, though it is a stringed instrument and it looks exactly like a small guitar, requires a different tablature than guitar. The strings are GCEA, not EADGBE, which makes the fingerings of the chords completely different.

Ukulele Makers, fiddle and uke playing robot

There's something about the guitar that makes it easy to seem like you know what you're doing, even if you don't. And the same goes for the uke, even moreso (simply because it has less strings). The uke rose to popularity only recently, and therefore after the rise of the internet and the death of hard-won knowledge (if I had the internet of today when I was learning guitar twenty years ago, I would never have tried to figure out songs by myself, and I would have understood much less of music theory because of it.) Players learning the uke today (and looking for or potentially uploading tabs), then would be less likely to have to learn as much to get by as a guitar-player of twenty years ago.

So the ukulele is somewhat easier to play due to both it's lesser number of strings, but also because of the means with which to learn new songs on it (which expands playing-availability to a wider audience of non-musically-trained players).

Finally, the guitar is a more widespread instrument than the uke, which brings with it more players who might know what they're doing, and tabulate and upload songs for others to learn from.

I speculate that these things make it less likely that a uke-player would be as equipped to decode songs as compared to a guitar player, and that this leads to less songs being tabulated and uploaded. Pound-for-pound, the numbers would be way off, but taken as a ratio of (real) uploaders to people searching for uploads, the uke ratio must look way different than that of the guitar. Also, for some reason, I say that a guitar player would be less likely to run a simple search rather than going to a trusted source (due to being more of a professional player? serious speculation here based on loose ideas of the comparative profiles of guitar-vs-uke players, I understand).

Mike and Jarvis' reggae-playing Ukulego robots 

Overall, when looking at the cyber-uke-sphere, it just smells like fertile soil for a place like a tab-generating robot to entice hapless players.

And though I may not know music very well, and I don't know how to actually program a robot, it can't be that hard to make a botsite that restates your search via a songtitle-corrector, a lyric-matcher, and a cache of chord names and respective key groupings.

It's just too bad they can't figure out how to actually decode the songs for us instead of just pretending to do it.

One day, Leonard B. Meyers will be proud...
Music, the Arts and Ideas, Leonard B. Meyers, 1967: Music as a Learned Probability System


Sandra -psychosandra- Holmbom

POST-POST SCRIPT
...something else about the chronologically stipulated evolution of the instruments respective to that of the internet...kind of like what happened to the ampersand in English vs. French, but not really.

Technologically-mediated cultural artifacts of both the ampersand and tab-generator software, see below.
The Ampersand
October 2012

Sunday, December 23, 2012

Culture as Learned Probability System

Music, the Arts and Ideas
Leonard B. Meyer, U. of Chicago, 1967

CULTURE
“It is impossible to stand outside of culture, for the models and categories we use on conceptualizing and ordering the world are necessarily limited to, if not determined by, those which are provided by our particular culture.” (viii)

“A culture, like a musical style, is a learned probability system.” (p17, footnote 21)


MEANING
“Meaning is when stimulus does not fit expectation.”

“Meaning is the difference between expectation and actual.”

Meaning, being also a measure of the uncertainty between antecedent and consequent is also related to probability.

PROBABILITY SYSTEMS
“Once a musical style has become part of the habit responses of composers, performers, and practical listeners it may be regarded as a complex system of probabilities. That musical styles are internalized probability systems is demonstrated by the rules of musical grammar and syntax found in textbooks on harmony, counterpoint, and theory in general.”
In the Tonal Harmony of Western Music, the tonic chord is
-most often followed by the dominant
-frequently by the subdominant
-sometimes by the subdominant
In Counterpoint, after a large melodic skip, the melody
-usually moves in the opposite direction, filling in the tones passed over.

DEVIATION
occurs by:
Delay – of expectation/norm
Antecedent Ambiguity/Uncertainty – equally probably constituents may be envisaged
Unexpected/Improbable – within the specific context
-deviation requires meaning, or more specifically: active creation of new meaning, which occurs within the temporal/tonal expectations


INFORMATION
Information is measured by the randomness of the choices possible in a given situation. If a situation is highly organized and the possible consequents in the pattern process have a high degree of probability, the information (or entropy) is low.

If, however, the situation is characterized by a high degree of shuffledness so that the consequents are more or less equi-probable, then information (or entropy) is said to be high.

Perception:
“What we perceive as the present is the vivid fringe of memory tinged with anticipation.”
-A.N. Whitehead, The Concept of Nature, in G.J. Whitrow, The Natural Philosophy of Time, p83, in Meyers, p89.

“The practically cognized present is no knife-edge, but a saddle-back, with a certain breadth of its own in which we sit perched, and from which we look in two directions in time.”
-William James, Principles of Psychology, 1950, p609

Predictions create the future:
“Prediction itself – that is, belief about the probable future – may make an incalculable difference in it.”
-Herbert J. Muller, “Misuses of the Past”, p12, Horizon 1 (March 1959), on Meyer, p90


The Determinism vs. Free Will Problem
Might it not be that the problem stems from a failure to distinguish between the relatively undetermined, non-statistical processes characteristic of the level of individual choice and those forces which, more determined and quasi-statistical, are operative on the higher levels of sociocultural history. It may be that the problem of determinism vs. free will stems from a confusion of hierarchic levels of analysis. (p97)

SATURDAY, OCTOBER 6, 2012
TUESDAY, DECEMBER 4, 2012
TUESDAY, JULY 10, 2012

Tuesday, November 27, 2012

The Anatomy of a Joke


Describing Meaning and Information
as a Function of Probability

-what makes a joke funny

“Meaning is when stimulus does not fit expectation.”
–Leonard B. Meyers

The funniness of a joke exists within an antecedent-consequent relationship, where expectedness of the consequent (the punchline) is plotted against both the ambiguity/clarity of the antecedent, and the time delay between the antecedent and the consequent.

The longer one waits to hear the punchline, the more possible punchlines one predicts. The longer one waits, then, the higher the probability that they will already have predicted the punchline, or something similar, rendering the joke less effective.

The more unexpected the punchline, the less the probability that one will guess it beforehand. This distance between what we expect, and what we get is what we might call the funniness of the joke.

Note: One can heighten the tension, or draw contrast by modulating the clarity or ambiguity of the antecedent, but this aspect of the dynamics of meaning will differ in a system of iteration, such as knock-knock jokes, or ‘inside’ jokes (internet memes/macro image series’), where one has “heard this joke before”.

For example:
“What happened when Chuck Norris got stabbed in a dark alley?” [antecedent]
“He punched his assailant in the face.” [consequent, highly probable]
--this is a reasonable outcome; it could have been predicted easily
“The knife bled to death.” [consequent, less probable]
--this is unexpected
--note the difference in effect, however, if one is already familiar with Chuck Norris jokes, as it changes the probability of prediction.

-something music 


IN MUSIC, a system of repetition, the interplay between antecedent and consequent is much more developed, but still derives the intensity of its meaning from the modulation of ambiguity/clarity of the antecedent, expectedness of the consequent, and the time delay between the two. Music can act as a very accessible entrypoint into thinking about information and meaning.

-something information network interaction
source: 
Ebon Fisher, Media Rituals, 1990-2000?

FROM INFORMATION,

Meaning is the difference between expectation and actual.”
(Meyers, p9-10)

“If we want to explain this absence of meaning [caused by the cognitive semiotic] and with it the emergence of a specifically human sense of reality, of time, space, and self, we have to assume that human cognition is based upon a very peculiar system of representation (or ‘pattern matching’) which allows us to process what is seen and heard at the same time, both in terms of stable patterns and of global, concrete and necessarily ‘fuzzy’ patterns. This double processing generates a difference between the stable pattern, which corresponds to what we have called memory, on the one hand, and an instable, always changing pattern corresponding to what one could name the ‘here and now’ or the ‘present’, on the other hand. In a conscious human mind the two patterns never merge completely.”
(Cognitive Semiotics, p122)

Deviation requires meaning, or more specifically, active creation of ‘new’ meaning.”
(Meyers, p9-10)

Aside: Between the passages immediately referenced above, we can begin to see the interplay of the past, the present and the future, and how they work to construct our world. Out of the double-processing of the past and the present, we create the present; and out of the simultaneous processing of all possible futures, we create the future. I would interject here furthermore, for there must be something said about the individual vs. the collective. The individual has the most power over the present, and the least over the future. The collective, though it conditions individual behavior (absolutely, though?) and helps to solidify the past (without any say from an independent self?), is largely more responsible for creating the future. Individuals cancel each other as they move forward, each on their own way. You can create your present, you can modify your past, but you cannot create your future, unless, of course, you are not.

TO ENTROPY,

Information is measured by the randomness of the choices possible in a given situation.
[does this read as ‘lack of connections’ between the choices?]

If a situation is highly organized and the possible consequents in the pattern process have a high degree of probability, the information (or entropy) is low.
[can this read as a ‘dense network’?]

If, however, the situation is characterized by a high degree of shuffledness so that the consequents are more of less equi-probably, then information (or entropy) is high.
[what is the relationship between network density and entropy?]

(Meyers, p11)

EPILOGUE,

“Both are looking at the world, and what they look at has not changed. But in some area, they see different things and they see them in different relations to one another. […] The transition between competing paradigms [^expectations] cannot be made one step at a time [via the left brain], but forced by logic and neutral experience. Like the gestalt switch, it must occur all-at-once [the right brain] or not at all.”
(Kuhn, p150), (“all-at-once”: Mass Transference Device, p50, p74)

NOTES:

Music, the Arts and Ideas
Leonard B. Meyers, 1967

Cognitive Semiotics
“Dealing with Difference: From cognition to semiotic cognition”, Barend van Heusden
Issue 4 (Spring 2009), pp. 116–132

The Structure of Scientific Revolutions
Thomas S. Kuhn, 1962

Mass Transference Device
2012

Bursts
A.L. Barabasi, 2010