Monday, November 20, 2017

THE END OF THEORY..Hold on, not so fast!





If it's possible to find the manifesto of "data science," the article published in far 2008 can be it. The main idea of research is that the classical scientific approach presented in the scheme form: a hypothesis model - the experiment, faced serious problems today. 


One of them - impossibility to check the theory experimentally. For example, it is impossible to prove or disprove the M-theory in quantum physics because the humanity has no sufficient resources yet to do the necessary experiments.


Approaching a similar problem, on the other hand, is possible using the data science tools. Instead of being focused on relationships of cause and effect, data scientists suggest to study the correlation between objects. Presence of large volume of data, and opportunities to analyze it, does the unnecessary existence of any theory or model in general!
It can be explained on the example of an algorithm of ranging of search pages in google. To define what of pages is more relevant deliveries to the computer not necessary to carry out the in-depth semantic analysis, and it is enough to be focused on statistics of attendance of a particular resource. The algorithm assumes that to us not important the reason motivating people to arrive in one way or another, to us it is important to trace and classify final behavior. With enough data, the numbers speak for themselves.

Search engines, e-marketing not the only scopes of the analysis of big data. Today data science is systematically integrated everywhere, since scientific community, and further into all spheres of human activity.


In fact, It's not about "the end of the theoretical study." Of course, we will need a theory in future. Theory guides us when we have limited data and gives us a place of solid ground to have discussions on. Moreover, a data science can supplement a classical science way of exploring a world around us. For example, results of machine learning process can point scientists where they should focus their activity and what model is meaningless.
The theoretical approach is not only a way of exploring, but it's also a way to structure and save the knowledge.
If we came to the start point of some new area where we don't have enough data the ML tools will be ineffective as well as statistical analysis. In this case, we need to develop our knowledge through building a sharp model, upgrading it and then, when we'll have enough data to analyze, we can use data science. Let's look at Ballistics. In case of data science, we'll need to do a large number of experiments to calculate a prediction for flying shell. And every time we when we want to calculate the trajectory, we need to spend a lot of resources. Or we can just use a one-line formula for every situation. It's all about a balance between approaches. Data screen helps theoretic to focus on usefull data, as well as theory can help to upgrade the data science methods.


The bottom line is we can't fully avoid a theory. It must be a part of the research process. The point is - the theory is just a first stage of data mining. Now we discovered a method which can cumulate and upgrade the theoretical approach - the data science


If to judge by the volume of development of studying of data, then it is possible to assume that shortly (even if not now), data science methods will occupy important weight in processes of decision-making in a scientific and business community.


If to whom it is interesting, with pleasure I will listen to criticism and possible shortcomings of a similar approach. 
You will find the reference to an article below. 

Tuesday, November 14, 2017

Goodhart's Law versus market speculations

"The more any quantitative social indicator is used for social decision-making, the more subject it will be to corruption pressures and the more apt it will be to distort and corrupt the social processes it is intended to monitor."

This is a variation of one common low in economics, known as "Goodhart's Law." A principle lies under it can be used to describe a fundamental low of market movement: When one man had found a pattern in market time series and started to use it, the patter disappears. It means that as soon as a trader start using the pattern, he becomes a part of the factors that affect the sequence. The result, the all previous dependences changes to a new one, that not necessary will follow the old pattern. 

Here is  an excellent visualization on how a "Goodhart's Law" works to a market patterns breaks:


Every time when a trader finds a new pattern and start using it, the pattern disappears. It can be variable from one to another: some models exist longer than average. Mostly because they are hardly recognizable. It is where a large hedge funds makes their money. They try to find a new pattern using more complicated models than concurrent. It ends when a crowd sees a pattern and start using it for speculations. By a crowd weigh a pattern is be broken and all starts again. Over and over.   

The similar effect I saw when my boss decided to use KPI(Key performance index) to evaluate the quality of work. It looks quite a safety until the primary measure of KPI correctly relevant to the real outcome of the job done. In my case, when KPY had started tracking the amount of "activity" inside corporative web portal and use it to evaluate the performance of the worker, lot of people start posting "flood" comments to increase those score. By the end of the story, the fraction of a useful information decreased dramatically. The recalculation of the KPY index only helped to solve the issue. 




Back to the data science problem, It is essential to find and use data that are not influenceable by the end user. In case of employer KPI, it must be a hard manipulated value, or some really valuable measure, for example, amount of sales for sellers. Otherwise, it can be turned into targets and lead to opposite effects one want to create. 
Time since that I have started look not only at a quality of data but assume how future events can affect already used information. It's fear of financial sector in particular, where human behavior changes market models constantly.

If you are interesting more about "Goodhart's low" you should check the following links: Wiki, DataSceptic podcast, and ribbonfarm.com


All best!




Friday, November 10, 2017

An ethical dilemma with hacked data

In addition to above-mention big data problem, I want to share with you one interesting article. In short, one company was trying to collect data for their research for a while and couldn't do it due to the technical complexity. But then some happened and database been hacked and published for everyone.  The published dump has been contained a mix of private and public data itself.  So, here is a dilemma: "Can this database dump now be used?", "Does it become "public"? or "whether to use this dataset to produce a socially useful research?"

The bottom line is the company didn't use the hacked data. They provide the list of arguments which I share. Here is some:

  1. Researchers have a limited capability to distinguish between public and private information within the hacked data. 
  2. May see private data when cleaning the data.
  3. Perhaps legitimizing criminal activity. 
  4. Violating users’ expectation of privacy. 
  5. Using people’s data without consent. 
  6. We want this data, but we don’t need it. Other data can be ethically collected and used
The only benefit of using the illegal information is a "faith in goodness powers of the research for men." But, honestly, it's a bull shit. The majority of research has a primary goal to increase the revenue of the company. The dirty pool game can break the fair concurrence in data-providers business. As a result, fewer companies will care about data security what can badly affect to the end user.



Hack attack on large credit company Equifax

In this case, the "black market of data" can occur. If a company needs some "sensitive" data for those research, they can just commission a hacking this data with the following publication. The company will wash hands of an affair shifting the blame on "a bad hacker." This kind of practice will finally remove borders in privacy.

Specialists of The University of Michigan comment:
"When using the hacked data, you reward criminal activity, and in this way, criminals will be motivated to find more ways to hack data. It is like buying a stolen bike from a criminal. Besides, the private data can come in (more) wrong hands so the private data will be spread more and more among more and more people. And because researchers have a limited capability to distinguish between public and private information within the hacked data, they may use private data or spread private data by accident. All the above will lead to a higher possibility of abuse of private data."


As a data science still a very young an against to journalism and don't have bases such as code of conduct, we need to be more careful in making decisions about what passes and what won't. It can be very complicated based on the fact that we can't evaluate an impact correctly for both cases. Let's say, we collect data for cancer research. For more performance, we need more information to mine. The results of our study would have a significant impact on man, sure enough. BUT, we can't calculate even closely the risk of concentration a massive amount of private, sensitive data in one place. If this kind of data would be used in bad faith, some the story can change unpredictably, and we faced the much worse questions. 

Sunday, July 2, 2017

Bitcoin price with google trends

A long ago I want to experiment and deal with a question of how Google trends can be implemented to analyze the financial markets. Here the opportunity has just turned up.


The main idea consists in the following: As the price is functioning of supply and demand, increase in demand for an asset will cause the growth of the price. The decrease in demand, or increase in the supply, respectively attracts reduction of the price. The assumption which I will check in at research is that the statistics of search queries on "hot trends" can correlate and advance price dynamics of a relevant "hot" asset.


In the beginning, I have decided to observe BITCOIN cryptocurrency. Today he more than "hot" because of the media interest and big volatility in a price.


We will need the following tools:


1. Data from the Google trend.


Google trends tools - the excellent example of a BIG DATA implementation. It shows dynamics of the popularity of a particular search query in time. Also, service provides tools for the analysis changes in inquiry and allows to compare keywords among themselves.


2. Historical quotes of BTC/USD


For that end, it can be used the quotas export from the MT4 terminal which is provided by the BTCe exchange.


After export of all data and formatting, I applied both lines in one chart. Here is the following graph:




The blue line is BTC/USD price, orange - the frequency of "bitcoin" query in Google search engine.


Already at this stage it evident that there is the correlation between both lines and dependence can be calculated. Here are the outputs:




We see that the correlation is 80.74%. It is evident that there are connections.


Unfortunately, Google doesn't provide detailed statistics for the entire period, so it can’t be calculated more even. With this data, there is a high error of approximation. Most precisely the dependence is reflected by exponential and polynomial regression:






Unfortunately, no conclusions at this stage can be drawn. While the data volume is extremely deficient, it can’t be the forecasting tool. Of Course, I need to consider the broader array of factors. For example, the decision of the Central Bank of Japan about legalization cryptocurrency increased interest to Bitcoin, which affect the price of BTCUSD.


At the following stage, I will try to find the big database which can be applied to the analysis effectively. I’ll see if they're available open sources solution, such as Quandl, for example.


Thursday, June 15, 2017

Privacy is a new luxury

"Privacy is a new luxury" - It seems that quite so it is possible to characterize our century of information.
I have come across one scandal with Sberbank and one fast-food restraint. Briefly: Sberbank has been charged with sale of the history of transactions of users in order of targeting advertising on the internet. Actually, it is Violation of Bank secrecy. And it is only one of a set of the cases recorded in recent years.At the same time, the companies even aren't protected, being covered with the "depersonalized" data and full "confidentiality."
Recently I have attended the conference on BIG DATA organized by Beeline. The main agenda of this meeting was just monetization of the clients given to activity (the home Internet, television, geolocation, etc.).
All this works by one principle: on the client the detailed statistics of the fact that he watches what sites he visits what purchases he makes In WHAT PLACES he is and AS it is FREQUENT collects. Further, these data are profiled and on sale to the advertising companies.
As a result, we receive:
I went to McDonald's - Receive advertising of the burger in the browser!
You look series on TV much - Receive cashback at a subscription to ivi
Sounds, it seems, harmlessly. And is further what? And also - algorithmization and the predictive analysis of your activity. Already now many recruiting agencies with results of Big data research. Ethics questions fade into the background here.
Of course, there are both pluses and apparent benefits for society. Some will consider the interest that thanks to technologies the necessary goods will easier get to the most needing consumers.
Personally my opinion such: while there is no big button "right to oblivion" - it is possible to be covered with "good intentions" as much as long.


Saturday, June 10, 2017

Russan ruble and oil price

Everyone knows how the impact what oil has on the Russian economy. But how it can be explained in the context of math and how it changes in time? I calculated the regression equation and pairwise comparisons for BR futures and USDRUB. All data I took from FINAM open database.


The first thing I did was import data arrays over a period of the 2012-2017 year.After a simple transformation of data the following
the chart has turned out:




On a vertical axis is a price of USDRUB, on a horizontal axis - BRENT future price. Even on this step, we see high dependence.
Output calculations will be the following:




The regression equation is:




With the main coefficients:




Variable A in the red square is approximation error.We see it goes least at exponential regression. Hereafter we’ll use this particular kind of regression. Here are the output charts:




Blue line -  line regression, Red- exponential.


Chart also shows the exponential dependence between USDRUB and BRENT price over the five years distances.
During that period share of oil and gas incomes in Russian trade balance have been about 47,6% (statistic average).Currently, the percentage has been declined to the smaller amount.


The table below shows how oil and gas share have been changing during the last decade.




Let’s explore how the correlation between USDRUB and BRENT was changing at this period. Likely that dependence should decline with reduct considering the part of oil and gas share in the trade balance. Here what show statistical calculations:


We should keep in mind that Central Bank of Russia moved to floating rate policy towards national currency in November of 2017. It reflected in investigating dependence. Remind that Central Bank of Russia had controlled ruble’s price by сorridor rule.


It shows how correlation had changed after removal the corridor rule. Best seen displays 2013 and 2014 year chart. Since 2014 USDRUB/BRENT have begun to obey the power low and we can see how correlation increased compared to previous periods.




2013 year


2014 year


This kind of dependence is observing up to the 2017 year. In sum, this gives a high statistical significance for prediction power of this model. Here is the equal for 2016:




However, there is one question left: Why the correlation has been reminding at the high level during the 2012 year? I’ll gonna think about it in upcoming research.


If calculates the theoretical price of Russian ruble with this formula, the following value will be around $64.99 (BRENT = 48.99). Considering that USDRUB is equal 57.04, we get the 14% error.It’s more than enough to say that similar ceased to reflect the reality.Also, the chart for 2017 shows the same:


BR/USDRUB 2017


It is evident that power dependence disappeared utterly. The correlation falls to 37.89%, which is the lowest since 2012. Why does it happen? It’ll be more clear at the end of 2017 when I can calculate the total data.In any case, it’s clear that the Russian economy is changing, as reflected in the graphs.


The main conclusion I drow from this research is that in the long term period there is a high power dependence between the USDRUB and oil price. This instrument can become useful for macroeconomic forecast inside trading strategies.


In fact, I found more questions than answers, so the investigation keeps up.Next time I’ll take a more significant time period with other macroeconomic indicators.
For those purposes, I need the stronger tool than excel.I think about R. let's see.

Sunday, May 14, 2017

About oil price

Ooh, long ago here I added nothing. Alas, until recently there was at all no opportunity to be engaged in independent researchers. Now it became slightly simpler with it, and it means that I will shortly publish some practices. Plus still is an idea entirely to move to the independent website shortly. Amicably, it would be necessary to give some comment on the global markets from the equipment or macroeconomic, especially against the background of the arriving news. It is remarkable that the other day Google recorded a historical maximum by requests for World War III. News of this sort always pushes people to invest in the "protected" assets and commodity. On the one hand, it looks quite reasonable, at the conflicting demand for raw materials will increase. But there is one problem: historically the prices don't keep long at the high levels, and correction will take away finally all collected profit. I have shown to one client who has wanted to invest for a long time in oil I the following chart:

The schedule shows dynamics of the price of oil from 1861 to 2011. At the same time, the blue line  price in nominal dollars, and the red line - in brought on inflation since 2011.


What can draw a conclusion? And very simple: Adjusted for inflation, the average price of oil of the hysteric woman was always not more expensive than $40 for the barrel. Any "carrying out" above finally was corrected. It means that oil purchase - initially unprofitable investment. The cost of providing a position, inflation and percent finally will destroy all profit, even if the price of some time grows.


Unfortunately, not so many people adhere to similar logic. And most of my clients don't consider similar historical extrapolation a sufficient argument and continue to play "random walks."

University Towns and Recession risk

The time has come for me to start looking for new apartments in the US. The logical question has appeared: What is the best area to re...