Showing posts with label dirty data. Show all posts
Showing posts with label dirty data. Show all posts

Tuesday, April 22, 2025

Is It Data

 
Surely every branch of science is special, and important, but some are more prone to credibility problems than others. Nutrition is one of them. Drug-based preventative health interventions are another. One suffers from almost totally unreliable survey reporting about what people eat, i.e., dirty data. The other suffers from a money problem, where money and power can make the data do whatever it wants, i.e., corruption. This is not to say that some branches of science are worthless, or absolutely untrustworthy, only that some need a stronger dose of skepticism than others. 

Researchers propose novel model to screen misreporting in dietary surveys
Jan 2025, phys.org

They found that 48% of food intake records in NHANES and 54% in NDNS had unrealistically low levels of energy intake [calories].

"This new model suggests that we should throw out large amounts of data, and nutritionists using dietary instruments may be unwilling to do that. However, continuing on just publishing erroneous data because it is too painful to acknowledge it's flawed, probably isn't the best way forward for nutrition science. I think as we go forward into the future many widely held beliefs that have been based on these problematical methods will need to be revised."

via Shenzhen Institutes of Advanced Technology of the Chinese Academy of Sciences: Rania Bajunaid et al, Predictive equation derived from 6,497 doubly labelled water measurements enables the detection of erroneous self-reported energy intake, Nature Food (2025). DOI: 10.1038/s43016-024-01089-5

Also: R. James Stubbs et al, Predictive equation helps estimate misreporting of energy intakes in dietary surveys, Nature Food (2025). DOI: 10.1038/s43016-024-01090-y , doi.org/10.1038/s43016-024-01090-

Image credit: AI Art - Supplements aka Small Pills - 2025


Experts challenge aspirin guidelines based on their undue reliance on a flawed trial
Apr 2025, phys.org

The American Heart Association (AHA)/American College of Cardiology (ACC) guidelines restricted aspirin to patients under 70, and more recently, the United States Preventive Services Task Force restricted aspirin use to patients under 60.

Researchers from Florida Atlantic University's Schmidt College of Medicine, and other distinguished collaborators, believe that both the AHA/ACC Task Force and the U.S. Preventive Services Task Force were unduly influenced by the uninformative, not null, results of the Aspirin in Reducing Events in the Elderly (ASPREE) trial. Specifically, this trial did not provide reliable evidence that aspirin showed no benefit in the age groups they enrolled.

"Absence of evidence does not equate to evidence of absence of effect."

via Florida Atlantic University: Janet Wittes et al, Aspirin in primary prevention: Undue reliance on an uninformative trial led to misinformed clinical guidelines, Clinical Trials (2025). DOI: 10.1177/17407745251324866


Tuesday, May 9, 2023

The Big Hard


Drinking alcohol brings no health benefits, study finds
Apr 2023, phys.org

You've been told that one drink a day is actually good for your health -- but it's so hard to believe, right? It's good for your heart! It can't be true right?

No, it's not true. We screwed up, for decades, using one bad study after another to support a crazy idea. Can it be true that not one person stopped and said, wait, that sounds too good, let me double check that study. Not until now. 

It's been named "former-drinker bias", and it will be in every public health textbook for the rest of time starting now, as an example of what can go wrong with biostatistics and epidemiological research. 

We heard rumblings of this a while back...

And now this article sums it up pretty well:
  • Former drinkers aren't lifetime abstainers -- For example, many studies tend to place former drinkers in the same group as lifetime abstainers, referring to them all as "non-drinkers," Stockwell said.
  • But former drinkers typically have given up or cut down on alcohol because of health problems, Stockwell said. The new analysis found that former drinkers actually have a 22% higher risk of death compared to abstainers.
  • Their presence in the "non-drinker" group biases the results, creating the illusion that light daily drinking is healthy, Stockwell said.
  • It's called "former-drinker bias"; and the reason it's been hiding in our public health research for decades? 
  • "This is an overview of a lot of really bad studies," Stockwell said. "There's a lot of confounding and bias in these studies, and our analysis illustrates that."

via Canadian Institute for Substance Use Research at the University of Victoria in British Columbia:  Jinhui Zhao et al, Association Between Daily Alcohol Intake and Risk of All-Cause Mortality, JAMA Network Open (2023). DOI: 10.1001/jamanetworkopen.2023.6185



Post Script:
Continuum of Risk
  • 2 standard drinks or less a week -- You are likely to avoid alcohol-related consequences for yourself or others at this level.
  • 3 to 6 standard drinks a week -- Your risk of developing several types of cancer, including breast and colon cancer, increases at this level.
  • 7 standard drinks or more a week -- Your risk of heart disease or stroke increases significantly at this level.

Bonus:
Partially unrelated, but still a good example of why science is hard:
HUGO (Human Genome Organisation) Gene Nomenclature Committee (HGNC), the body that names genes, has changed 27 genes to avoid being confused by Excel's default naming protocols.

For example, SEPT2 is the short name of a gene called Septin 2....
-Scientists rename human genes to stop Microsoft Excel from misreading them as dates
Aug 2020, The Verge

Poisoning the Well



I almost feel irresponsible for posting this picture above, so a mandatory public service announcement is in order: Do not put bleach in your air vents.

And now for something totally different:

Two types of dataset poisoning attacks that can corrupt AI system results
Mar 2023, phys.org

The researchers began by noting that ownership of URLs on the Internet often expire—including those that have been used as sources by AI systems. That leaves them available for purchase by nefarious types looking to disrupt AI systems. If such URLs are purchased and are then used to create websites with false information, the AI system will add that information to its knowledge bank just as easily as it will true information—and that will lead to the AI system producing less then desirable results.

The research team calls this type of attack split view poisoning. 

There is another way that AI systems could be subverted—by manipulating data in well known data repositories such as Wikipedia. [This has been a tactic by authoritarian governments since its inception.]

via Google, ETH Zurich, NVIDIA and Robust Intelligence: Nicholas Carlini et al, Poisoning Web-Scale Training Datasets is Practical, arXiv (2023). DOI: 10.48550/arxiv.2302.10149

Post Script:
Another public service announcement -- the datasets used by today's deep learning artificial intelligence are not stored locally, they are stored as URLs which have to be accessed at the time of execution.

In other words, the Stable Diffusion LAION dataset is not a bunch of pictures; instead, it's a bunch of url's of pictures, like a url with ".jpg" at the end. This is good because it makes the memory storage for 5 billion images much smaller, because you're storing the link to the picture, not the actual picture. 

For anyone who's done anything on the internet for more than 5 years, you know what link rot is, and why it should make you really confused as to how people think the current crop of AI magic will continue to work as all the urls rot out, making the dataset smaller and smaller, and the quality of the output worse and worse. 

(And this isn't even considering the intentional data poisoning attacks described above.)

Also, double check the thumbnail for this post, which is an ad-poisoning injection about "this one trick" to get the dust out of your air vents by pouring bleach in them, itself designed not to advertise a product, but simply to get you to click so the broker can charge both parties for clickthroughs, even though no "eyeballs" took place, and as much as this shouldn't be happening, it is, and it's now in The Big Dataset in the Sky, poisoning our artificial intelligent systems. 

Wednesday, September 7, 2022

Weaponized Delivery Packages for Misinformation


One day customers will only want to do business with those who harvest their data sustainably.

Twitter pays $150M fine for using two-factor login details to target ads
May 2022, Ars Technica

"As the complaint notes, Twitter obtained data from users on the pretext of harnessing it for security purposes but then ended up also using the data to target users with ads," Federal Trade Commission Chair Lina Khan said. "This practice affected more than 140 million Twitter users, while boosting Twitter's primary source of revenue."


Feds seize SSNDOB marketplace that listed personal data of 24 million people
Jun 2022, Ars Technica

Social Security Number Date of Birth (SSNDOB) like the walmart of personal data.

More fallout from the Chainalysis revelation, which is basically that every transaction you make on the blockchain is public knowledge, so with some good network software, you can track people and money like a first grade math problem. 

Further readings:
Inside the Bitcoin Bust That Took Down the Web’s Biggest Child Abuse Site
Apr 2022, Mike McQuade, WIRED [soft paywall]


Facebook is receiving sensitive medical information from hospital websites
Jun 2022, The Markup via Ars Technica

Experts say some hospitals’ use of an ad tracking tool may violate a federal law protecting health information. (You don't say)

A tracking tool installed on many hospitals’ websites has been collecting patients’ sensitive health information — including details about their medical conditions, prescriptions, and doctor’s appointments — and sending it to Facebook.

The Markup tested the websites of Newsweek’s top 100 hospitals in America. On 33 of them we found the tracker, called the Meta Pixel, sending Facebook a packet of data whenever a person clicked a button to schedule a doctor’s appointment. 

Clicking the “Schedule Online” button, filling in the booking form, or clicking the “Finish Booking” button on a doctor’s page sent the following information:
  • text of the button clicked
  • doctor’s name
  • doctor's field of medicine
  • search term used to find doctor: “pregnancy termination"
  • condition selected from dropdown menu: “Alzheimer’s”
  • first name
  • last name
  • email address
  • phone number
  • zip code
  • city of residence  entered into the booking form,
  • names of patients’ medications
  • descriptions of their allergic reactions
  • upcoming doctor’s appointments
  • name and dosage of a medication in our health record
  • notes we had entered about the prescription
  • response to a question about sexual orientation
The Markup also found the Meta Pixel installed inside the password-protected patient portals of seven health systems. 
You heard the man; this is a stick up. 

Technical sidenote:
“The evil genius of Facebook’s system is they create this little piece of code [the pixel] that does the snooping for them and then they just put it out into the universe and Facebook can try to claim plausible deniability,” said Alan Butler, executive director of the Electronic Privacy Information Center. “The fact that this is out there in the wild on the websites of hospitals is evidence of how broken the rules are.” (So the pixel is like a dematerialized AirPod?)

Further reading on body brokers and biodata:
Sapiens For Sale, Network Address, Aug 2022

Network structure of Agents in Tsuchiyu Onsen, Tohoku University, 2022


Kochava faces legal action over sale of location data
Aug 2022, BBC News

The company, founded in 2011, says on its website that it "complies with all user data privacy and consent regulations".

And they do, because there aren't any.


Data privacy bill would give you more control over info collected about you
Aug 2022, The Conversation via Ars Technica

"Excludes deidentified data"
-American Data and Privacy Protection Act (Frank Pallone, hello NJ)

How hard is it to 'de-anonymize' cellphone data? 

Not hard:

We study fifteen months of human mobility data for one and a half million individuals and find that human mobility traces are highly unique. In fact, in a dataset where the location of an individual is specified hourly and with a spatial resolution equal to that given by the carrier's antennas, four spatio-temporal points are enough to uniquely identify 95% of the individuals.
-Unique in the Crowd: The privacy bounds of human mobility. Yves-Alexandre de Montjoye et al. Sci Rep 3, 1376 (2013). https://doi.org/10.1038/srep01376

Four datapoints. That's 2013 by the way.

Post Script:
Here's the thing: we've all watched the promise of tech and the internet curdle into (at best) invasive, advertisement-saturated, rent-seeking bullshit and/or (at worst) weaponized delivery packages for misinformation, bigotry, and occasional incitements to genocide and violence. I think we're all reaching our saturation limit for being monetized, marketed to, invasively tracked, and charged a premium for devices and services that enable those things. I think we're looking for relief from all that, not variety of opportunities to experience it. -Snark128, "Meta sparks anger by charging for VR apps", Financial Times via Ars Technica, Jun 2022 https://www.ft.com/content/e8910bad-b873-407d-b1ca-46eb4ceb3db2 


Thursday, August 11, 2022

Greatest Retronym in History


When it comes to AI, can we ditch the datasets?
Mar 2022, phys.org

Synthetic fucking data. They're making synthetic data to train the robots. And that makes us analog data. Me and you, our faces, our fingerprints, our gaits, gestures, voices (and most especially our consumer behaviors), are analog, starting now.

First there was the acoustic guitar, then dairy milk ffs. Hopefully, when we finally cede control to the omnibot envelope, we don't go the way of the flip-phone. 

via MIT: Paper: Generative models as a data source for multiview representation learning. openreview.net/pdf?id=qhAeZjs7dCL


Physiological signals could be the key to 'emotionally intelligent' AI, scientists say
Apr 2022, phys.org

You got any more of that analog data?
They're coming for your sweat, your biodata. You are the training set for the artificial humans of the future. 

via Japan Advanced Institute of Science and Technology JAIST: Shun Katada et al, Effects of Physiological Signals in Different Types of Multimodal Sentiment Estimation, IEEE Transactions on Affective Computing (2022). DOI: 10.1109/TAFFC.2022.3155604

Image credit: Jared Michael

Friday, April 29, 2022

Baselines, Biases and Big Data Problems - The Secrets of Statistics and the Magic of Metrology



The paradox of big data spoils vaccination surveys
Dec 2021, phys.org

Image credit: Labyrinth by Liqen, Miami 2011

So guess what -- big data is really good at minimizing sample size errors, but also at magnifying systematic biases such as nonresponse bias like how vaccinated people are more likely to respond and marginalized groups less so -- in other words, bad data. Yet because the sample size is so large, we are tricked even more into thinking it must be good data. "Biases in the data get worse with bigger sample size." -phys.org

Here's an example: "Two in 10 respondents did not have a college degree, compared with four in 10 of all U.S. adults—and race and ethnicity—the fraction of black and Asian respondents was only half of what it is in the general population."

"Worse than no survey at all" they say.

Thanks Facebook, but you can keep your invasive mass surveillance system of the entire American population to yourself. 

via Harvard: Seth Flaxman, Unrepresentative big surveys significantly overestimate US vaccine uptake, Nature (2021). DOI: 10.1038/s41586-021-04198-4


On research methods and data quality:

Meng said he began thinking about the problems posed by big data during a visit to Harvard a decade ago by a U.S. Census Bureau official. The official met with a group of statisticians and asked them about the handling of data sets that were becoming available covering large percentages of the U.S. population. Using the hypothetical example of tax data collected by the IRS, he asked whether the statisticians would prefer a sample covering 5 percent of the population that they knew was representative of the larger population or IRS data that they weren't sure was representative but covered 80 percent of the population. The statisticians chose the 5 percent. "What if it was 90 percent?" the Census Bureau official asked. The statisticians still chose the 5 percent, because if they understood the data, their answer would likely be more accurate than even a much larger set with unknown biases.

And now for the hard stuff:
Completely unrelated image; I'm just collecting pictures from science articles of people holding vials in their fingers: Wastewater Filtration, Pacific Northwest National Laboratory and Andrea Starr, 2022

Drinking alcohol to stay healthy? That might not work, says new study
Nov 2021, phys.org

This is such a great example of how statistics works (and how it doesn't). The key word is "baseline".

We've been told for years, forever?, that one glass of red wine a day is not only not-bad, it's actually good for you. In health-speak, they call that "protective".

But this study shows us that there is no group of people that we can use as a baseline, who ---doesn't--- drink and yet who also ---doesn't--- have other health issues caused by having a history of substance abuse. In other words, when the only people who don't drink are people who are recovering from being addicted to alcohol (gross exaggeration), it means you can't find a baseline for a "normal" person. And if you can't find a baseline, then you can't measure anything. 

The majority of the alcohol abstainers at baseline were former alcohol consumers and had risk factors that increased the likelihood of early death. Former alcohol use disorders, risky alcohol drinking, ever having smoked tobacco daily, and fair to poor health were associated with early death among alcohol abstainers. Those without an obvious history of these risk factors had a life expectancy similar to that of low to moderate alcohol consumers. The findings speak against recommendations to drink alcohol for health reasons.

I tell my friends this story, and they ask the first question, a good question -- what about people like Seventh Day Adventists, or Muslims, they don't drink, why can't we use them? And the answer is the reason why health science is hard.

The reason you can't use groups of people who don't drink, is because that group would likely not represent the much larger group of, let's say, all the people in America. It doesn't even matter if it's a small group; making it bigger won't help. It doesn't work because it doesn't match. You can't compare the two groups because the people aren't the same. 

Another way we see this, and one which is becoming more evident to those who can fix it, is how certain groups of people (like undocumented immigrants) are under-represented in the data. If the majority of the datapoints are White, Christian and middle class, and you're none of those, then it's possible that the data is not relevant to you. Your baseline isn't represented, so whatever health effects you're trying to measure, they aren't being compared to someone like you. 

via Public Library of Science: John U, Rumpf H-J, Hanke M, Meyer C (2021) Alcohol abstinence and mortality in a general population sample of adults in Germany: A cohort study. PLoS Med 18(11): e1003819. doi.org/10.1371/journal.pmed.1003819

Partially Related:
This is also a correlate to the story of how lead was discovered in the air -- a geochemist was trying to do such sensitive work (to measure the age of the Earth!) that he kept picking up extra lead in his results, and could not figure out where it was coming from. But it turns out that it was in the air, having been vaporized in the internal combustion engines in our cars. And that story implies that until then, all other experiments being done were "wrong" because they didn't exclude the excess lead from the otherwise "normal" background. 

Last One -- Palmar Sweating:
During a nuclear war scare (1950's), all experiments into palmar sweating at a research institute had to be abandoned because the base level of the response had become so abnormal that the tests would have been meaningless.  (p188)
-The Naked Ape, Desmond Morris, 1967

Monday, March 14, 2022

Look Mom No Data


AKA From Deep Learning to Deep Reasoning

DRNets can solve Sudoku, speed scientific discovery
Sep 2021, phys.org

You can teach a machine to recognize a dog by showing it 1,000 pictures of dogs, Gomes said, but scientific discovery is not like that.

"You are not going to have lots and lots of labeled data," she said. "And in general, the examples you have are not exactly what you are looking for, but then you reason about what you know scientifically about the domain, and you can infer new knowledge."

Key to DRNets is the idea of an "interpretable latent space." Basically, it gives DRNets the ability to reason about the constraints of the domain—in this case materials science—from input data.

They started with Sudoku -- de-mixing overlapping handwritten Sudoku puzzles—grids. The computer had to separate the puzzles into two solved Sudokus, without any training data, which it was able to achieve with close to 100% accuracy.

The researchers then put DRNets to work on a real-world problem: automating crystal-structure phase mapping of solar-fuels materials, using X-ray diffraction (XRD) patterns. Crystal-structure phase mapping involves separating the source XRD signals of the desired crystal structures from "noisy" mixtures of XRD patterns, a task for which labeled training data are typically not available. ... DRNets was able to identify and separate a total of 13 crystal phases (single-phase materials) in 19 unique mixtures of the single-phase materials. ... DRNets' findings, verified using manual analysis, enable the discovery of complex mixtures of crystalline materials that convert solar energy into storable solar chemical fuels.

via Cornell University: Di Chen et al, Automating crystal-structure phase mapping by combining deep learning with constraint reasoning, Nature Machine Intelligence (2021). DOI: 10.1038/s42256-021-00384-1