Hi Cl?lia,
first of all, i would like to congratulate you for this research that you have presented here! I have read your pdf two/three times, and i would like to make some comments to it.
I have also done some works like that in astrology (animodar and a statistical study of religiosity), and i also know some of the pitfalls we encounter. Also, statistics is a tool that i use a lot in my work (computer science). However, please note that the comprehension of someone outside the study is very limited, so some of my comments may be out of sync with your work..
Here it goes:
1) In your first page, you talk about the size of the test set. Basically, as you can't study the all population of suicides, you must select a representative set (randomly selected), which you can somehow say it represents the (bigger) one. The idea behind is "Confidence interval", ie, you want to measure how confident you are about the smaller set being representative of the entire population. The article in wikipedia explains very well the subject (
http://en.wikipedia.org/wiki/Confidence_interval), and it has a practical example also. Basically, for an interval of confidence of about 95%, you should have a random set of about 800 cases (it is an equation).. Polls for government elections have sets of this size, so it is pretty standard..
2) On page 3 you say "Many astrologers question whether astrology is adequate for statistical studies, and I am beginning to agree with them: statistics is the best way to set apart individual factors, however it misses complex configurations.". However, i believe you're wrong: astrology is not only adequate for statistical studies, but it is build over statistics. Let me clarify: when someone says that "a person that has a dignified Venus is fond about artistic subjects", it is in fact saying that "there is a probability of X% that a person that has a dignified Venus is fond about artistic subjects". It is of great importance to know that value of X.
Also, in my opinion, complex configurations aren't missed by statistics, but by people, ie, complex configurations leads to excess of information, and so, begins to be hard to work on.
In my work of the animodar, and even more on my (unpublished but live presented) work about Religion, i've processed about 200 test cases, and many configurations (like if the ruler of house 9, house9, Jupiter, and Saturn had any configuration (aspect, disposition, same ruler, etc.) with the same thing for house 1, and other conditions). Of course, i used skyPlux to compute all this data in 2 seconds (yes, 2 seconds but many hours to do the programming), generated a table with hundreds of columns, and then used a software to "extract" automatically patterns from the data (Rapid Miner at
http://rapid-i.com/). The hard work was then for interpreting the rules and graphics that it generated. Just to say that after some things, you have to automatize things, or you will get lost. However, it is not statistics that it is of no use, but it is data overrun..
You give an example of complexity by Ptolomy, but also that example can be read that "there is a probability of X% that if each of the lights should be pivotal and each of the malefics [respectively] should be present or else in opposition...". It is just a (complex) rule that can be evaluated as true or false..
3) On page 2 you talk about Venus having malefic aspects with the ruler of House 8, twice as often than in the control group. Why did you reject this only because it is too simple? You then say that your data set is too small, but 320 cases is big enough for a semi-conclusion. Or you tested this with less cases? You also talk about differences of 102% and 109%. What those numbers really mean, i didn't understand.. My max is 100%!
4) On page 8 you say that you've discovered examples about the use of the part of spirit in your data set but also in the control set. It would be good to present some numbers for comparing your data set with the control set..
5) On page 9, did you test any of those aphorisms? Number for comparing the probability of occurrence in your data set against your control set?
6) On page 26 you present your conclusions, like "the Moon frequently gets into configuration with malefics or the ruler of the 8th House. The expected configuration is not by aspect, but by rulership instead". Again, you should have numbers to compare both sets. Maybe something that you seemed to notice more is not really occurring more, but stayed in your mind, and it is a kind of a bias. The same thing for the other aspects of your conclusion.
7) Having the previous things more concrete, you could present your findings more clearly. For instance, you say that the moon aspects malefics or the 8th ruler, or that mars appears angular in difficult houses, and sometimes in conjunction with mercury. However, we don't get to understand how much importance each of those statements have. Maybe the moon aspecting malefics accounts for 70% of the suicides. But as you don't show any numbers, the statements tend to get generic, and we can't extract the most importants from there..
8 ) In page 30 you show a guide that lists the importance of each statement in the previous point. Again, it would be important to know the numbers that back you up..
Bottom line, i say again that you have an excellent work here, and the only thing keeping it from being perfect, in my opinion, is the lack of numbers like i've mentioned. They are important to show us the weight that each sentence of the conclusion has. For instance, if a sentence has a probability of 90% it is pretty strong, but 10% is almost nothing. For instance, if i say:
- The moon aspecting malefics accounts for 90% of the suicides is different than saying
- The moon aspecting malefics accounts for 10% of the suicides.
The second is almost of no importance..
I hope you keep up the good work, and don't mind the fact of me saying things like "you should this", "you should that", etc. Accept my apologies for that, but it is that you have done a great and important work here..
Thanks,
Jo?o Ventura