Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Are there really algorithms for judging the truth? I think the closest we can get is heuristics for identifying whether or not something is being actively deceptive, but it's a big epistemological leap to say that we can definitely determine truth in an algorithmic way.

Google is very good in these situations, but Google has a hard time differentiating between a popular movement and a single determined astroturfer. I'm sure Google would rather display a commitment to good journalism ("You want to know about topic X? Here's what people are saying about it.") rather than a pursuit of the truth, at least when it comes to organic results.



Sure, if the thing to be judged has a limited number of inputs and factors and outcomes...which is something that does not apply to the content of the Web, of course.

But in other cases where the answer is basically, "yes" or "no"...if an algorithm chooses the right answer, doesn't that count as truth? Take for example, the Apgar score, which consists of giving a newborn a score of 0 to 2 for five categories: complexion, pulse rate, reflex, muscle tone, and breathing. http://en.wikipedia.org/wiki/Apgar_score

A baby that scores 8 is considered to be doing well. A baby with just 5 merits additional examination.

This simple numerical score is credited with revolutionizing obstetric medicine and is still used today, even though it was first published in 1953. http://www.newyorker.com/archive/2006/10/09/061009fa_fact?cu...

I guess it's debatable whether this counts as an "algorithm"...but it involves the mechanical evaluation of a "score", with set guidelines for when such score is dangerously low. The fact that the scoring is done through human judgment is peripheral to the matter...it's not hard to imagine a machine being able to perform the same test (perhaps more accurately) were it cost-practical to build such a machine.

One more interesting note: The Apgar score was not developed by an expert. It was developed by a female doctor who had never delivered a baby (nor had a baby) herself, but felt that doctors were giving up too quickly on sick babies.


Yeah, that's more of a heuristic for determining the health of an infant. An algorithmic approach would give you a definite yes-or-no for all baby inputs.

And the analogy breaks down here because the online ecosystem adapts very fast. It would work if there was some sort of gene that led to premature births plus high Apgar scores (the equivalent of a malicious website that ranks well).

That's why search is so tough: search engines and gaming have coevolved to the point that in plenty of areas, search quality would go down if people weren't trying to game the algorithm.


Heuristics are algorithms and algorithms don't have to give definite answers.


No, that's just not true. A heuristic is a procedure that gives no guarantee of a correct answer. An algorithm is a deterministic procedure that does guarantee a correct answer. Heuristics are very useful in making difficult decisions that don't have definite answers, but they are not algorithms.


Erm. Algorithms aren't necessarily deterministic. http://en.wikipedia.org/wiki/Nondeterministic_algorithm

Primality testing is a very good example of a useful algorithm that gives an answer that is probably correct.


Thank you for the correction. I'd been using terminology ineffectively. The way I was taught it was:

- An algorithm is a step-by-step process for transforming any valid input into a specific valid output. - A heuristic is a process that brings you closer to something that looks like an answer, but does not necessarily find the best possible answer.


I've been thinking a lot about this problem lately.

Perhaps in the era of the web, the right answer is to have a kind of error range for sourcing.

If you see the same fact quoted in three places on the web, you track that there are three instances, and the number of independent sources is anywhere from 1-3.

On the web, reshares and linkings far outnumber independent synoptic views of the same data. So in a way it's more accurate to take the opposite approach, assuming it's a single source.

You start upping the number of independent sources when you can prove it. Like, there are multiple camera views. Or, you somehow verify that the independent observers are real people, and you make some best guess about whether their accounts are independent (that's trickier though).




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: