Monday, November 20, 2017

THE END OF THEORY..Hold on, not so fast!





If it's possible to find the manifesto of "data science," the article published in far 2008 can be it. The main idea of research is that the classical scientific approach presented in the scheme form: a hypothesis model - the experiment, faced serious problems today. 


One of them - impossibility to check the theory experimentally. For example, it is impossible to prove or disprove the M-theory in quantum physics because the humanity has no sufficient resources yet to do the necessary experiments.


Approaching a similar problem, on the other hand, is possible using the data science tools. Instead of being focused on relationships of cause and effect, data scientists suggest to study the correlation between objects. Presence of large volume of data, and opportunities to analyze it, does the unnecessary existence of any theory or model in general!
It can be explained on the example of an algorithm of ranging of search pages in google. To define what of pages is more relevant deliveries to the computer not necessary to carry out the in-depth semantic analysis, and it is enough to be focused on statistics of attendance of a particular resource. The algorithm assumes that to us not important the reason motivating people to arrive in one way or another, to us it is important to trace and classify final behavior. With enough data, the numbers speak for themselves.

Search engines, e-marketing not the only scopes of the analysis of big data. Today data science is systematically integrated everywhere, since scientific community, and further into all spheres of human activity.


In fact, It's not about "the end of the theoretical study." Of course, we will need a theory in future. Theory guides us when we have limited data and gives us a place of solid ground to have discussions on. Moreover, a data science can supplement a classical science way of exploring a world around us. For example, results of machine learning process can point scientists where they should focus their activity and what model is meaningless.
The theoretical approach is not only a way of exploring, but it's also a way to structure and save the knowledge.
If we came to the start point of some new area where we don't have enough data the ML tools will be ineffective as well as statistical analysis. In this case, we need to develop our knowledge through building a sharp model, upgrading it and then, when we'll have enough data to analyze, we can use data science. Let's look at Ballistics. In case of data science, we'll need to do a large number of experiments to calculate a prediction for flying shell. And every time we when we want to calculate the trajectory, we need to spend a lot of resources. Or we can just use a one-line formula for every situation. It's all about a balance between approaches. Data screen helps theoretic to focus on usefull data, as well as theory can help to upgrade the data science methods.


The bottom line is we can't fully avoid a theory. It must be a part of the research process. The point is - the theory is just a first stage of data mining. Now we discovered a method which can cumulate and upgrade the theoretical approach - the data science


If to judge by the volume of development of studying of data, then it is possible to assume that shortly (even if not now), data science methods will occupy important weight in processes of decision-making in a scientific and business community.


If to whom it is interesting, with pleasure I will listen to criticism and possible shortcomings of a similar approach. 
You will find the reference to an article below. 

Tuesday, November 14, 2017

Goodhart's Law versus market speculations

"The more any quantitative social indicator is used for social decision-making, the more subject it will be to corruption pressures and the more apt it will be to distort and corrupt the social processes it is intended to monitor."

This is a variation of one common low in economics, known as "Goodhart's Law." A principle lies under it can be used to describe a fundamental low of market movement: When one man had found a pattern in market time series and started to use it, the patter disappears. It means that as soon as a trader start using the pattern, he becomes a part of the factors that affect the sequence. The result, the all previous dependences changes to a new one, that not necessary will follow the old pattern. 

Here is  an excellent visualization on how a "Goodhart's Law" works to a market patterns breaks:


Every time when a trader finds a new pattern and start using it, the pattern disappears. It can be variable from one to another: some models exist longer than average. Mostly because they are hardly recognizable. It is where a large hedge funds makes their money. They try to find a new pattern using more complicated models than concurrent. It ends when a crowd sees a pattern and start using it for speculations. By a crowd weigh a pattern is be broken and all starts again. Over and over.   

The similar effect I saw when my boss decided to use KPI(Key performance index) to evaluate the quality of work. It looks quite a safety until the primary measure of KPI correctly relevant to the real outcome of the job done. In my case, when KPY had started tracking the amount of "activity" inside corporative web portal and use it to evaluate the performance of the worker, lot of people start posting "flood" comments to increase those score. By the end of the story, the fraction of a useful information decreased dramatically. The recalculation of the KPY index only helped to solve the issue. 




Back to the data science problem, It is essential to find and use data that are not influenceable by the end user. In case of employer KPI, it must be a hard manipulated value, or some really valuable measure, for example, amount of sales for sellers. Otherwise, it can be turned into targets and lead to opposite effects one want to create. 
Time since that I have started look not only at a quality of data but assume how future events can affect already used information. It's fear of financial sector in particular, where human behavior changes market models constantly.

If you are interesting more about "Goodhart's low" you should check the following links: Wiki, DataSceptic podcast, and ribbonfarm.com


All best!




Friday, November 10, 2017

An ethical dilemma with hacked data

In addition to above-mention big data problem, I want to share with you one interesting article. In short, one company was trying to collect data for their research for a while and couldn't do it due to the technical complexity. But then some happened and database been hacked and published for everyone.  The published dump has been contained a mix of private and public data itself.  So, here is a dilemma: "Can this database dump now be used?", "Does it become "public"? or "whether to use this dataset to produce a socially useful research?"

The bottom line is the company didn't use the hacked data. They provide the list of arguments which I share. Here is some:

  1. Researchers have a limited capability to distinguish between public and private information within the hacked data. 
  2. May see private data when cleaning the data.
  3. Perhaps legitimizing criminal activity. 
  4. Violating users’ expectation of privacy. 
  5. Using people’s data without consent. 
  6. We want this data, but we don’t need it. Other data can be ethically collected and used
The only benefit of using the illegal information is a "faith in goodness powers of the research for men." But, honestly, it's a bull shit. The majority of research has a primary goal to increase the revenue of the company. The dirty pool game can break the fair concurrence in data-providers business. As a result, fewer companies will care about data security what can badly affect to the end user.



Hack attack on large credit company Equifax

In this case, the "black market of data" can occur. If a company needs some "sensitive" data for those research, they can just commission a hacking this data with the following publication. The company will wash hands of an affair shifting the blame on "a bad hacker." This kind of practice will finally remove borders in privacy.

Specialists of The University of Michigan comment:
"When using the hacked data, you reward criminal activity, and in this way, criminals will be motivated to find more ways to hack data. It is like buying a stolen bike from a criminal. Besides, the private data can come in (more) wrong hands so the private data will be spread more and more among more and more people. And because researchers have a limited capability to distinguish between public and private information within the hacked data, they may use private data or spread private data by accident. All the above will lead to a higher possibility of abuse of private data."


As a data science still a very young an against to journalism and don't have bases such as code of conduct, we need to be more careful in making decisions about what passes and what won't. It can be very complicated based on the fact that we can't evaluate an impact correctly for both cases. Let's say, we collect data for cancer research. For more performance, we need more information to mine. The results of our study would have a significant impact on man, sure enough. BUT, we can't calculate even closely the risk of concentration a massive amount of private, sensitive data in one place. If this kind of data would be used in bad faith, some the story can change unpredictably, and we faced the much worse questions. 

University Towns and Recession risk

The time has come for me to start looking for new apartments in the US. The logical question has appeared: What is the best area to re...