Wednesday, August 3, 2016

Preventing Spam Account Registration

(v0.2, updated 10/5/2016)

I am interested in the topic of preventing spam account registration. Therefore, I have collected, organized and commented some resources from the Web. The current post discusses various approaches at a very high level. I plan to expand the depth and breadth of this post in the future. Any comments are highly appreciated!

---------

We define a spam account of a Web service as an account that is not for legitimate usage of the service. Rather, the purposes could be:

  • increase follower count, view count, etc.
  • manipulating voting
  • sending advertisements (pharmaceutical)
  • spreading malware
  • SEO
Within the complex cyber criminal ecosystem, some people are specialized at creating spam accounts and sell them. Just search "buy facebook account" on Google and you should find sellers like:

  • https://buyaccs.com/en/
  • http://www.allpvastore.com/

Apparently, many activities supported spam accounts could have very negative impact on the Internet. So our goal is to detect and stop spam account registration, or delete spam accounts as soon as possible. However, completely preventing spam account registration might be too difficult, so we can also have a softer goal: to increase the cost of spam account registration as much as possible.

Many approaches have been proposed to mitigate spam account registration. Below are some of them.


1. Preventing automatic account registration

Nothing is more convenient for the attacker if he can register a large number of spam account automatically, either using a single machine or a group of bots. To prevent such attack, the defender needs to differentiate between human visitors and bot visitors. Some approaches are:

1.1 Captcha

The assumption is that visual recognition is an easy task for human, but a hard task for bots. However, the visual recognition challenges can be cracked by bots [1]. Also, even if the challenge cannot be cracked by computer for now, the challenge can still be solved by human crowdsourcing.

1.2 Registration randomization

The HTML code of the registration Web page or even the registration process can be randomized. It probably does not affect a human user, but might disrupt bots that needs to parse the page and follow the process. Some discussion about this idea can be found in [6].


2. Requiring additional information during registration

This type of approaches requires the attacker to provide additional information, such as an email account, a phone number, etc. If obtaining such information costs more than verifying the information, the defender can then apply them to increase the cost of attackers.

2.1 Mobile number and/or email account

This type of information is easier to verify, typically by a confirmation message. It is still possible that the attacker can obtain a large number of mobile numbers or email accounts. For example, Gmail allows you to have (infinite) variations based on one email address. Some services also sells mobile numbers for receiving SMS. But the cost of registration spam accounts is increased.

In addition, the defender can occasionally ask users to reconfirm their mobile numbers or email accounts. This further increases the difficulty of using spam accounts, because the mobile numbers or emails used at the registration time might already become unavailable (e.g., sold to someone else or get blocked) [7].

2.2 Personal identification information

The defender can requests real name plus some personal identifiers (PID) such as passport ID, SSN, etc. Then the defender checks whether the PIDs matches the name, or whether the PIDs are consistent (e.g. used by the SSA website). However, in many cases the defender needs the help of a third party, typically an government agency to do the verification. In addition, some users might not like submitting such sensitive information.


3. IP rate-limiting and blocking

Number of account registration allowed per IP address can be limited. In addition, if a large and abnormal amount of registration comes from the same IP address, the IP address can be blocked [5].

However, this approach might cause collateral damages to benign users because:
  • Dynamic allocation of IP addresses [3]: the blocked IP address might be used later by a benign user.
  • Proxy, gateway: all users behind such a device are blocked from registration.
  • Tor [4]
In addition, this approach might be bypassed if the attacker use proxy, etc.


4. Detecting Spam Accounts

4.1  At Registration

The defender has different kinds of information about a new registration request:
  • Account information, including username, phone number, email etc.
  • Behavior information, such as the time gaps between two consecutive operations.
  • Low-level information, such as IP address, user-agent, etc.

The defender can then train a detector that can separate spam account registration from normal account registration. Such a detector can be constructed based on human insights. For example, [7] proposes to use regex to capture patterns of spam account names. They have patterns because many of them are generated automatically based on some rules. So the defender can sort of reverse engineer the rules.

Of course, the defender can directly train a machine learning model, if labeled training data are available. The defender can also do unsupervised learning (e.g. clustering) to get more insights about spam accounts.

4.2 After Registration

After an account is registered, the service will have more information to differentiate spam accounts and normal accounts. Example signals are:

  • Activities done after registration.
  • Characteristics of the friendship network

These signals are likely to be helpful for a detector. However, at the time of detection, the spam account might already have been used in malicious activities.

4.3 Challenges
  • For a service with large number of users, even a tiny rate of false positive (e.g. 0.1%) will cause huge collateral damage, and thus is not acceptable.
  • The model might become less accurate overtime, due to concept drift.
  • The attacker might maliciously change its behavior to evade the learning algorithm. Or the attacker could even pollute the model.  Several reports have shown that many machine learning algorithms are not robust under adversarial scenarios.


5. Post-registration tracking

The service can store cookies at the client side. Then if the same cookie is used by multiple accounts, there is a possibility that the client registered spam accounts. The service can consider using very persistent cookies, such as the evercookie, which works even cross-browser.


6. Collaborating with other services

6.1 OAuth

Why not offload the nasty task to services with spam account registration prevention, such as Google and Facebook [2]?

6.2 Threat database

For example, the service can check an IP against known IP blacklist.


References:

[1] Snapchat's Latest Security Feature Defeated in 30 Minutes
http://www.infosecurity-magazine.com/news/snapchats-latest-security-feature-defeated-in-30/

[2] http://stackoverflow.com/questions/170152/prevent-users-from-starting-multiple-accounts

[3] https://www.quora.com/How-does-an-ISP-assign-IP-addresses-to-home-users

[4] http://www.computerworlduk.com/tutorial/security/tor-enterprise-2016-blocking-malware-darknet-use-rogue-nodes-3633907/

[5] https://en.wikipedia.org/wiki/IP_address_blocking

[6] ShapeShifter: The emperor’s new web security technology | The Good, The Bad and the Insecure
https://blog.securitee.org/?p=309

[7] Thomas, Kurt, et al. "Trafficking Fraudulent Accounts: The Role of the Underground Market in Twitter Spam and Abuse." USENIX Security 2013

Tuesday, November 10, 2015

Reflections on a Data Scientist Job

(v0.1 - 2015/11/10)

I've been working as a data scientist intern for several months now. Here, I want to summarize a little bit on my experience, and also list several things that I found important. Hopefully, I will achieve these goals soon:)

I think there are several tasks for data scientists, including analysis, research, building internal data products and building external data products. The goal for analysis is mainly answer some direct business and engineering questions. A data scientist can expect to receive various questions from different teams constantly. Research tries to tackle more difficult and complex questions in order to provide deeper insights. Then, some of the analysis or research results could be turned into internal data products, mostly as internal web services. Thus, other members in the company can reuse the analysis for their tasks. Finally, if the methodology is reliable and the result is good, we can release the data service as part of the company's product to the outside world!

In a large company, each person can specialize in a particular area. While in a small company, the data scientist typically needs to be good at everything (full snack?). But I think no matter the size of the company, the following goals are always important:

1. Code Quality. Many people think that the data code only needs to run once, so one can just write it in an ad hoc way. But this is not true based on my experience. The code will be reused and the project will become more and more complex. So all the software engineering wisdom needs to be applied in data code as well. Some extra investment in the beginning will save significant amount of work later on.

2. Use Established Tools. One way to produce high quality code and save time is to use established tools. This both includes development tools (e.g. IPython) and data science libraries (e.g. pandas).

3. Communication. Before working on a problem, it would be important to clearly understand the purpose and desired outcome. A lot of the time, the requester might not fully understand the complexity of the data. So some simple measures, such as mean, might not work as expected. After the job is done, it is also important to convey the result to other people. The write up, in the form of an email or slides, needs to be accurate and concise.

4. Feedback. Actively seek feedback from different types of people. Engineers, sales, customer managers, etc., all give very good and diverse comments. The comments could help you improve the workflow and get new ideas. If working on some data products, create an early demo and seek early feedback.

5. Learning. There are several aspects related to learning. First, one shall learn about what's going on inside the company in order to actively find new opportunities of applying data analysis. Some folks might not realize that their tasks can be helped by a data scientist. Second, one shall improve its skills, including statistics, machine learning, etc. Third, one need to search for related work of the current project. Some data tasks are unique, but many others are similar. So learning from other people's experience can significantly facilitate one's own projects.

6. Planning. Apparently, there are many questions waiting to be answered, just like there are many development tickets to be addressed. My experience is that it is usually hard to finish all of them.  So one needs to prioritize. Also, in the beginning of a project, always start with a simple solution. A "bad habit" for people from academia is that they tend to make things over complex... Complexity is sometime required to get your paper published, but it is usually not a friend in a company.

I did not talk about big data, because the dataset I am focusing on is rather small. But again, I would say no matter the size of the data, these goals always applies. In addition, the trend of technology is hiding more details of lower levels, so that people can put more effort in building the tower of technology higher. That means in the future, the interface for analyzing small and big data would be pretty much similar.


Some interesting data science article and data blogs:

[1] The Top Mistakes Developers Make When Using Python for Big Data Analytics
https://www.airpair.com/python/posts/top-mistakes-python-big-data-analytics#3-mistake-2-not-tuning-for-performance

[2] Data Scientist: The Sexiest Job of the 21st Century
https://hbr.org/2012/10/data-scientist-the-sexiest-job-of-the-21st-century/

[3] Data Research from OkCupid
http://blog.okcupid.com


Saturday, September 12, 2015

Language and Intelligence




Part 1 International Students

I learned from my girlfriend and her friend that your language ability determines your intelligence, particularly for an international student. Intelligence can be decomposed into two parts. The first part is whether you can understand some materials and come up with good ideas, which I assume should not be difficult for many students. The second part is whether you can convey your thoughts to others (intelligence). Unfortunately, I found that many international students, including myself, are struggling, because our ability to speak the language (e.g. English), can sometime significantly limits how we express our understanding and ideas. This could give others an impression that this guy is not intelligent. Or think about an extreme case where I was thrown to Moscow. Without knowing a single word in Russian, I cannot express anything through language. So to the Russians, I am equivalent to a complete idiot.

So it is very important to improve one's language to improve the perception of one's language by others.

Part 2 Mathematics

Many people say math is hard, and they cannot learn math well because there are not that intelligent. But it might be the case that they have the intelligence to learn math, but they do not have the language to read math. Using the example in Part 1, if I am in Moscow, I cannot book a hotel or visit the hospital, although nobody would consider these two tasks are beyond a normal human being. I cannot do these things because I don't know Russian. Similarly, you cannot do math if you don't know the language. The language of math looks a little bit intimidating, but I think it should not be more difficult than a human language (might be related to Chomsky hierarchy [2]). So I think people who are afraid of learning math can just view it as learning a new language. Learning a new language requires constant input and practice, so this explains why we should learn things (e.g. math) constantly as well.


References:

[1] The Figure. http://edl.ecml.at/Portals/33/images/EDL_Logo1.jpg

[2] https://en.wikipedia.org/wiki/Chomsky_hierarchy






Thursday, March 26, 2015

Risk Analysis of the Flight 9525



The crash of Germanwings Flight 9525 is really a great tragedy. It is even more sad to remember that there were multiple large-scale aviation accidents, including MH370, MH17 and TNA222, in the past two years. Many people, including myself, would wish that we can have better technologies and policies to reduce the likelihood of such events or even prevent them completely.

In the ongoing investigation of the Flight 9525 accident, we learn that the co-pilot locked the cockpit while the caption was outside, and then brought down the plane. The exact motive of the co-pilot is unknown. Apparently, airlines have tests the mental conditions of pilots, but currently there is no report indicating that the co-pilot is abnormal [1].

In this article, I would like to discuss how we might be able to decrease the risk of such accident. I am inspired by the following article written by professor Juliette Kayyem [2]: Was 9/11 safety precaution a flaw? From the title, we can already know one main point of that article: the cockpit lock-up mechanism designed to prevent 9/11-style attack becomes a problem when one of the pilots goes wrong. The author has suggested to have an emergence password so that no one can block the access to the cockpit. I definitively think this is a good idea, but we should think deeper by considering the risks of different threats, and how these risks tangle together.

There are many threats to an airplane: hijacking, mechanical errors, pilot errors and malicious pilots. Each threat has a risk value which can be simply calculated as likelihood * impact. We will only focus on the likelihood part in this article. We want to reduce the likelihood of every threat to be lower than certain threshold. However, this case clearly shows the difficulty, because one mechanism that reduces the likelihood of a threat (e.g. hijacking in this case) could increase the likelihood of another threat (e.g. malicious pilot). In this particular case, the cockpit lock-up mechanism is not good because the likelihood of malicious pilot has been increased above the threshold.

One might think it is necessary to have the lock-up mechanism to defend against hijackers, and we have to sacrifice on other aspects. But I don't think so. I think there are many other ways to reduce the likelihood of hijacking, such as security check points, on-board security guards which I've seen in Chinese domestic flights several years ago, and background check of passengers. This line of defenses are probably able to reduce the likelihood of hijacking to an acceptable level. On the other hand, however, we do not have reliable methods to prevent malicious pilots. As we have discussed previously, mental tests are not useful in this case at least. And due to the complexity of this job, we have to give many authorities to the pilots. Being able to unlock the cockpit, therefore, become an important defense line for malicious pilots. But unfortunately, this defense line was turned off for Flight 9525...

Another idea is to consider self-flying airplanes. After all, we already have self-driving cars. At least, the airplane could become a remotely controlled drone in emergence. This would not only help Flight 9525, but other cases when the pilots lost conscious, such as the Helios Airways Flight 522. But having an self-flying system introduces new threats such as software bugs or even vulnerabilities, which are major threats of all kinds of digital systems now. Should we trust human or machine?



References:

[1] Lufthansa CEO: Germanwings copilot passed medical exams http://www.cnn.com/2015/03/26/europe/lufthansa-ceo-germanwings-crash/index.html

[2] http://www.cnn.com/2015/03/26/opinions/kayyem-germanwings-co-pilot/index.html


Sunday, February 15, 2015

The Path of Sergio Leone




(v0.1)

Sergio Leone is one of my favorite movie directors. One can definitely attribute his success to genius and hard working. However, I think it is also interesting to take a look at his path of making films, as we might be able to learn something from it.

He has roughly made 10 films over 25 years (1959 - 1984) [1]. His first two films were rather bad according to the rating on imdb. However, these two films probably gave him enough experience to make a better one. So he made the third one, a Fistful of Dollars, in 1964. This movie was a huge success, but with one problem: he basically plagiarized the story of Yojimbo, by Akira Kurosawa. Personally, I am fine with his deed, because he made a great film after all, and he has compensated Akira. But more importantly, I think this step might be necessary for a young director like him, as he lack the experience of writing a good story. Such imitation is probably the fastest way to become a master.

Since the 1964 movie was a huge success, the wise idea is to make sequels of it. It can hone his ability further and get reputation and money quicker, with very little risk. So we have For a Few Dollars More (1965) and The Good, the Bad and the Ugly (1966). Then, Leone was already a mature director, and it was the time to climb the high mountain in his life. The previous three movies are all Western, but the stories are constrained by each other. He needed to break out from the trilogy in order to fully exploit his creativity, while also utilize his experience in Western films. So he directed the Once Upon a Time in the West in 1968, which is one of the greatest movies of all time. This movie also made him one of the greatest directors.

The four films had probably exhausted his creativity in Western  settings, so his eye turned to the past of Mexico and produced Once Upon a Time... the Revolution in 1971. It is also a great movie, because at this time, Leone already reached the top level, so it was impossible to make a low quality one anymore.

Then in the next 13 years, he stopped directing. It probably because he was tired and needed some rest. Also, he was preparing the next big shot. The next movie, finally arrived at 1984, was in a completely different setting compared to all his previous films. It is the Once Upon a Time in America which describes the lives of several gangsters in the New York City. I personally think this movie has reached the apex of filmography. The story, the acting, the scenes, the music ... are all the best. He is a true master.

Then, of course, we just need to expect one masterpiece after another from him, until his death. However, the ending of his life came rather soon because his body cannot catch up with his great mind. He died at 1989 when preparing Leningrad: The 900 Days.

I think he had a fantastic life with invaluable contributions to the humanity. I also think his path is similar to many great minds in other fields, such as academia (e.g. replacing films with research publications in the main text). I hope we all can get some inspirations from the paths of these forebears.


Additional remarks:

  • We should also emphasize the contribution of Ennio Morricone, who made superb music for Leone's movies. Their life long collaboration is also worth remembering. 

References:

[1] http://en.wikipedia.org/wiki/Sergio_Leone#Filmography

Saturday, January 31, 2015

Notes on the GHOST Bug

The recent GHOST bug discovered in glibc is a heap buffer overflow that could potentially lead to arbitrary code execution. I am interested to learn about this bug because I am working on heap buffer overflow defense. So I read the post written by Qualys Security Advisory, which really provides excellent explanation of it! [1]

This article just contains some of my notes when learning this bug. Hope they will also be helpful to others and please feel free to provide your comments.

(1) Why can't we detect this bug earlier?

It has been said that this bug existed since 2000. So an important question is why can't we detect it earlier? The article written by Qualys indicates that they found it through a manual code review. So probably the code has not received enough eyeballs previously.

On the other hand, I also think this bug is fairly easy to be detected by fuzzing. Because it is actually very easy to create test inputs and the oracle in this case. On the other hand, it is probably not easy to find the Heartbleed vulnerability through fuzzing, because both the test inputs and the oracle are hard to build.

I have wrote the following simple program that could trigger the vulnerability. We can think it as a very simple fuzzer.

https://github.com/movingname/Toys/blob/master/C/GHOST2.c

We can use the AddressSanitizer as the oracle. So I used clang + AddressSanitizer to compile it. Then when I ran it, AddressSanitizer indeed reports a heap buffer overflow.


I guess one could do a round of fuzzing for all this kind of functions in libraries. Maybe more bugs can be found?


(2) procmail exploit

The article [1] shows how we can exploit this bug in procmail using

/usr/bin/procmail 'VERBOSE=on' 'COMSAT=@...ython -c "print '0' * $((0x500-16*1-2*4-1-4))"` < /dev/null

However, this command has some omissions (the ... in the middle). Actually, one can run

/usr/bin/procmail 'VERBOSE=on' 'COMSAT=@'`python -c "print '0' * $((0x500-16*1-2*4-1-4))"` < /dev/null

to trigger the glibc detection.

In addition, the length of the input is important. Apparently it cannot be too small. But it also cannot be too large because procmail will detect the overflow. there is a tiny window that will trigger the overflow.



References:

[1] http://www.openwall.com/lists/oss-security/2015/01/27/9

Thursday, January 15, 2015

Inflation and StackOverflow.com




(v0.1)

I really enjoy asking and answering questions on the StackOverflow.com, which has some exciting features. For example, it is mainly managed by the community, not the admin. So users on StackOverflow.com not only contributes information (Web 2.0), but also contributes "computation" (Web 3.0?). Another feature is the scoring system. A user can get scores based on votes of questions and answers she posts. This is a quite good incentive for users to offer there knowledge, because the score reflects one's progress and ability, and can be used in job interview. This second feature separates StackOverflow.com from traditional mail list based Q & A.


However, this reward mechanism also comes with an issue which I call it the disadvantage of late members. That is, it is more difficult for a late member to get the same amount of score than an early member. This is mainly because early questions are in general more important yet easier to answer than late questions. For example, user X asks how to sort a dictionary in Python. User Y easily answers this question. Since this question is quite common for new Python users, we can expect that many user will find this answer and give it a vote up. Y thus earns high score simply through a simple answer. Now after one year, Z enters StackOverflow.com. Z is much more knowledgeable in Python than Y, however, there are not such easy and rewarding questions for Z now. So Z's score might never go beyond Y's.

This issue could hurt the participation of such sites. New comers could lack strong incentive to make contributions, because they can hardly find questions to answer, and they cannot catch up with the early members anyway.

We can think about how to address this issue. I propose two simple ideas here. The first idea is to reduce the score of early members overtime. However, reducing one's "possession" sounds bad and might hurt the early members. A slightly different idea is to introduce inflation. That is, the Q & A site should increase the score for a vote up overtime, thus giving more scores to new members. Maybe the inflation in economy also serves similar purpose. After early some early people have accumulated a large amount of wealth, it would be hard for late people to catch up, because the rich people can have better return simply through safe investments such as government bond. Inflation, in some degree, could alleviate this issue.

Actually, the solution used by StackOverflow.com now is to have a recent ranking (e.g. in past year) together with overall ranking. The recent ranking would be a fairer play ground for everybody. However, the recent ranking is not as stable as the overall ranking, so I think the inflation idea still make sense. Of course, StackOverflow.com can complement the recent ranking with permanent badges (e.g. Top 1 in one month badge).

A related article:

Why I no longer contribute to Stackoverflow
http://michael.richter.name/blogs/why-i-no-longer-contribute-to-stackoverflow/


Saturday, November 22, 2014

A Tale of Crowdsourcing and Diversity




Crowdsourcing is a hot topic recently, and I think it is a promising paradigm for solving a lot of problems. Diversity is an important phenomenon in complex systems such as the human society, and I am very interested in understanding the nature of it. Some of my previous blog posts, such as The Amazing Diversity is dedicated to this topic. In addition, our recent paper on a vulnerability disclosure program is also inspired by these two keywords. But how are these two concepts connected?

In the summer, I have read a famous Chinese Wuxia book called the Ode to Gallantry (侠客行). I found that the tale in the book serves as a perfect example for understanding crowdsourcing and diversity. I will very briefly introduce the story here, and please stop reading if you don't want to see this spoiler.


Sometime during the ancient China, many kongfu masters will be hijacked to a mysterious island every 10 years by some mysterious guys. These masters never return. It turns out that two top kongfu masters have obtained an old martial art book with undeciphered text. They have tried hard to understand the meaning of the book, but failed. Therefore, they decide to invite (or hijack) kongfu masters and ask them to decipher it (crowdsourcing). These masters, unwilling to go to the remote island at first, will soon be attracted by the book and concentrate entirely on the decipher task. However, many years have passed and still no one has figured it out.

It is not surprise that the protagonist of this story solves the problem. The unique advantages of him are:
  • He is illiterate. Therefore, he tends to understand the writing as graphs.
  • He has seen a graph-based kongfu book before, and this further guides him to interpret the text as graph.
  • He has little knowledge of kongfu before, and therefore does not have much prejudices and biases (the Einstellung effect).
Then, the protagonist learns the super kongfu in the book and become invincible in the world.


We can see that the initial crowdsourcing effort fails because there is a bias. And this bias is overcame by increasing the diversity of the crowd. Here, the protagonist is drastically different from the rest and thus significantly increases the diversity of the pool. This is one reason why diversity is important to crowdsourcing.

In general, crowdsourcing is still in its infancy and we are still exploring the meaning of diversity. There are many questions to be answered.



Reference:

[1] The picture. http://pjh568.gotoip2.com/data/attachment/forum/201207/28/101235i5w4ir1hj4ik2g45.jpg






Sunday, November 2, 2014

On Academic Presentation



(v0.1)

I will present our paper at a CCS workshop next Friday. Then I will present my thesis proposal in the comprehensive exam next next Friday. Facing these two important occasions, I decide to summarize my current understanding on presentation. This is NOT a collection of advises, because I am far from a good academic speaker. I simply hope this article may raise some discussions and help you think about what will lead to a good academic presentation.

Here I have several points to share:

(1) A clear story flow in the presentation is of top priority. The flow can grasp the attention of the audience. As others have said [1], the flow is much more important in slides than in paper, because the audio channel is more brittle. In addition, a good flow will also help the presenter to remember what to say.

I think there are at least two types of flows:
  • The logic flow of research. The audiences should know the natural transition between research steps. Thus they will appreciate the work. 
  • The knowledge flow. We need to introduce enough background before going into details. Also, make sure that terms etc. are understandable.
In addition, try to only have one story line. It is true that a research project usually expands to several branches. But they will interrupt the flow and confuse the listeners.

(2) Presentation is a process of convincing others. The listeners will be convinced if the study is rigorous and the language is accurate. Do not over claim.

(3) Make the presentation tight. Try to connect things together. Try to refer back to previous important points. This actually improves the complexity of the presentation structure, and people enjoys complexity. Similar strategies have often being used in movies. Lock, Stock and Two Smoking Barrels is a perfect example.

(4) Presentation is also a form of teaching. Try to think what the audience will learn from it.

(5) We have our own styles in presentation. I feel it is in general hard to copy other's style. For example, native-English-speakers can talk about jokes and funny pictures (e.g. the one used in this blog), which are sometime hard to understand, not to mention to speak, by non-native speakers. Nonetheless, even without these funny elements one can still make a good talk. I sometime think too much "fun" will actually have negative effects, i.e.,  "amuse to death".

Here are general steps I take for preparing a presentation. Please feel free to comment on them and provide your own opinions:

(1). Have a rough story line first.

(2). Turn the story line into slides. Focus more on the completeness of the information.

(3). Practice lightly and then update the slides. At this stage don't expect them to be perfect. Also take a look at similar talks to "steal" good presentation ideas.

(4). Write the scripts for all slides. At least write outlines for each slide. You don't need to read them, but you need them to remind you about the story line. Also, written text is easy to be studied and improved.

(5). Practice seriously.

(6). Present to others. Your adviser or research collaborators are the best choices. They know your research, but they are not trapped by myriad of details like you. So they can give very good suggestions on improving the story line! People with enough knowledge background (e.g. your lab mates) are also good. They can tell you which part is unclear or confusing. Also, try to collect creative ideas of presentation from others.

(7). Improve slides, practice, improve slides, ....

In general, you will feel unconfident and uncomfortable in the beginning, because the quality of your taste is always ahead of the quality of your work [2]. However, as long as you keep improving it, the final version will be very good. Furthermore after a well preparation, you will not only have a great talk, but also find new research ideas!





References:

[1] 博士五年总结(三), http://blog.sina.com.cn/s/blog_946b64360101dych.html

[2] Ira Glass on Storytelling, http://vimeo.com/24715531

[3] The picture. http://assets.diylol.com/hfs/ae1/38e/525/resized/business-cat-meme-generator-boss-wished-me-luck-on-the-presentation-like-i-need-it-52c717.jpg

Saturday, October 11, 2014

English Name or Not?




(v0.1)

As a Chinese student in America, an important question to ask is: should I choose an English (first) name? Those who against this idea usually provide the following points:
  • The original name defines your identity.
  • You should respect the original name because it is given by your parents.
  • If I am good, others will correctly pronounce and remember my name anyway. 
Some of my American friends, Indian friends and Chinese friends are holding these points. Sometime ago, I've also watched a Youtube video in which an American student advocates these points to some Taiwan students.

Other people, such as Philip Guo [2], support the idea of choosing an English name when moving to an English-speaking country.

And here is my opinion: although I currently do not have an English name, I agree that Chinese students (or possibly other East Asian students) studying in America should find a English (first) name. Obviously, the English name is easier to pronounce and to remember by both the natives and students from other countries. The English name can also tell the person's gender, which in some situations are more convenient. I guess the reason that most Indian students don't choose an English name because their original name is relatively easy to pronounce and already tells the gender, at least based on my experience. After all, English and Hindi both belong to the family of Indo-European languages.

Furthermore, I disagree with the three points that are against finding an English name. To refute them, we can look at the opposite direction: what did some Westerners do when they were in China. During the Age of Discovery, many Jesuit priests came to China and played an important role in the communication between civilizations. These priests all used Chinese names, such as 利玛窦 (Matteo Ricci, the man in the above figure),汤若望,郎世宁, which are still known by many Chinese today.

Also, having a second name is actually a part of traditional Chinese culture. Ancient Chinese people use their style name (字), rather than their real name in the daily lives. And it is actually impolite to call one using the real name. It is not a bad idea to consider the English name as a style name.


References

[1] The picture, http://www.faculty.fairfield.edu/jmac/sj/scientists/riccimap.gif

[2] http://www.pgbovine.net/choosing-english-name.htm

Saturday, September 20, 2014

A Quick Analysis of Facebook Bug Bounty Program




(v2, updated 10/15/2014)

Nowadays, Web companies have been relying on vulnerability reward programs (VRP, also called bug bounty programs) to discover vulnerabilities in their products. Basically, a white hat (good hacker) can submit a vulnerability discovery report and then get some money back. We have written a preliminary paper analyzing a related program called Wooyun, and please take a look if you are in general interested in this new paradigm of improving security.

Facebook is one of the companies that embrace this idea, although Facebook is generous sometime (see this and this), :P. FB also hides information about what vulnerabilities have been discovered, or the details of each white hat's accomplishment (e.g. how many vulnerabilities one has discovered, and when). FB only provides a list of white hats who have contributed to Facebook security every year, at this page.

Anyway, we can start with this page and do some quick analysis. The data is obtained by 9/20/2014. First, there are 670 names on the list (there are several cases when multiple names appear in one line and separated by commas, and we will count each name alone). Quite a lot, isn't it? But it is possible that some enthusiastic white hats contributed every year and leave their name multiple times, so we also count the number of unique names, which is 516.

Next, we count the number of white hats each year, shown in the following table:

TimeWhite Hat Count
2014 (up to 9.20)191
2013255
2012126
201155
Prior to 201143

We clearly see the trend: more and more players are joining this game, and the number roughly doubles every year:) I guess VRP is really a promising idea (please see our paper for more discussions).

There is also an interesting fact: a lot of white hats are only active in one year. To show this, we create another table counting the white hats based on number of years being active:

Number Years being ActiveWhite Hat Count
1402
282
326
45
>=51

So far, there are 402 who have only appeared in one year's thank list. And we can see that the white hat count distribution is highly skewed. Much few white hats are active for more than one year. And there is only one person who has been thanked all the time! This probably shows that the value of this kind of VRP not only lies in a few experts, but also in a large number of people. But since we don't know how many vulnerabilities each white hat contributes and the severity of them, the conclusion is hard to make. Still, this observation is consistent with what we claim in our paper.

You might wonder who is the "all the time" person, and the answer is: Szymon Gruszecki. You can access his personal page here.

Please feel free to discuss by leaving a message. Thank you for your time!


Update:

Facebook has released some interesting statistics of its bounty program here:
https://www.facebook.com/notes/facebook-bug-bounty/bug-bounty-highlights-and-updates/818902394790655

Some interesting points:

  • From the statics we see that there is a huge number of invalid reports. The valid rate is only 4.7%. Why?
  • It says that "One of the most encouraging trends we've observed is that repeat submitters usually improve over time. It's not uncommon for a researcher who has submitted non-security or low-severity issues to later find valuable bugs that lead to higher rewards." Actually, we plan to investigate this issue further in our data set.
  • The country rank: Russia -> India -> USA -> Brazil  -> UK





References

[1] The picture. http://america.aljazeera.com/content/dam/ajam/images/shows/Real%20Money%20with%20Ali%20Velshi/SG_FB2_1460.jpg



Saturday, August 23, 2014

Bugs and Patches for Papers



In an earlier article, Writing Like Compiling, I have made some connections between programs and papers. This article makes a connection from a different perspective.

For publications in Computer Science (and possibly other domains), there is a problem. A paper could contain bugs: errors, unclear sentences, missing backgrounds, etc. These bugs might caused by the knowledge gap between the authors and the readers. Or they simply arise due to the conference-driven publication paradigm. Such paradigm puts researchers on a fast race and leave them less time to ponder and polish their work [2]. These bugs inflict readers minds and eats up their time. Some smart readers might find ways to fix those bugs, just like an advanced user finds a bug in a program and makes a patch for it. However, since there is not good way to share the fix, and a paper is usually fixed after the camera-ready version, this paper patches only stay in a paper copy as some red marks...

Programs, too, are not perfect after release. However, software developers and users will constantly discover new bugs and apply corresponding patches. And this model generally works well. After all, there seems to be no alternative way. Therefore, I think we need to treat papers as programs, and creates ways for reporting bugs and sharing patches. A first step is to store these paper bugs and patches in some database and enable readers to search for them. However, we want to avoid the detachment between the papers and the patches, so we could allow some energetic readers to fork a paper and make version 2.0 of that paper. This fork ability is very common in the opensource software community [3].

In general, isn't it a bit ironic that Computer Science, the field that aims to digitalize paper-based information, still record its cutting-edge findings in papers?


References

[1] The picture. http://www.rebeccaheflin.com/wordpress/wp-content/uploads/2013/08/rejected-writing.jpg

[2] Fortnow, Lance. "Viewpoint Time for computer science to grow up." Communications of the ACM 52.8 (2009): 33-35.

[3] Raymond, Eric. "The cathedral and the bazaar." Knowledge, Technology & Policy 12.3 (1999): 23-49.

Saturday, July 12, 2014

Two Downsides of Privacy



v0.2

These days people are all talking about privacy, partly thanks to Mr. Snowden's effort. While I definitely support the the individual right of privacy, I want to talk about two downsides of privacy here. But again, I strongly agree that each individual should have full control of her or his information. I just think sometime we might want to trade privacy for more important things.

The first downside is that privacy could lead to distrust. An often used example for supporting privacy is: you got drunk one night and shared a photo of drinking on Facebook. While your friends might like it, your boss doesn't. So people have been designing advanced access control technologies to keep your boss away from your "little secrets". Some people might even avoid using Facebook, considering that some companies require employees' Facebook passwords [2]. However, such personal information disclosure could help others understand you more and thus enhance the relationship. On the other hand, if a person cannot be traced at all on the Internet, he or she will be a mystery in others' eyes. And will you trust a mysterious figure?

The second downside is that privacy might obliviate a person. From East to West, from past to present, an eternal pursue of the mankind is immortality. At least, a person wants to leave something to this world after death, which could only be achieved by a small group of people in the past. This digital age enables immortality to everyone, in the sense that their words and activities on the public Internet can be recorded and kept almost forever. Search engine could retrieve one's words and return to somebody in the future, and we could imagine it as a kind of conversation between the dead and the live. However, privacy would make these information secret to only a few people or even one person, through technologies like encryption. If that person happened to pass away, then his words also gone with him, if no one else knows the password. What a pity if Einstein II encrypted his remarkable theory and then passed away accidentally. Would it be better to publish it on a blog, like this blog?

Update: we all feel really sad about the MH17 tragedy. It has been said that there are more than 100 AIDS researchers on board. And I hope we can rescue their ideas and thoughts as much as possible.



References:

[1] The picture. http://blog.static.abine.com/blog/wp-content/uploads/2011/10/privacy.jpg?e835a1

[2] http://www.usatoday.com/story/money/business/2014/01/10/facebook-passwords-employers/4327739/

Friday, June 27, 2014

Heartbleed and the Paradox of Security Professionals


The recent Heartbleed vulnerability of OpenSSL shaken the whole Internet. Yet such incident might not be a total surprise because OpenSSL only had one full time employee and received $2000 a year as a donation [1] before the incident. Such support is by no means enough for the developers to produce high quality code and test the software product comprehensively. I guess the even hackers who keep searching for vulnerabilities inside OpenSSL have much more funding.

But the situation is changed now, as tech giants agreed to fund OpenSSL for at least 3.9 millions in three years. Even a Chinese mobile company, Smartisan, announced a donation of 1 million yuan ($160,000) [2]. You might already feel the strangeness of this event: a mistake makes millions of money!

The more astonishing thought is that if OpenSSL developers did a better job by not introducing the vulnerability, then they will still starving and suffering in poverty! This paradox does not only apply to OpenSSL, but probably to every company that needs security. For such a company, if the security team is doing a good job, then the company's CEO might feel the security team is redundant because nothing bad happens. Although the CEO might not that dumb to fire the security team, nonetheless the CEO could not appreciate the effort of the security team and might not raise their salary. Thus for the security team, there is hardly any incentives to do better. Rather, they might just want to meet the minimum requirements, or even accept some security incidents to attract the attention from company managers.

Do you know how to break this paradox?




References:

[1] Tech giants, chastened by Heartbleed, finally agree to fund OpenSSL. http://arstechnica.com/information-technology/2014/04/tech-giants-chastened-by-heartbleed-finally-agree-to-fund-openssl/

[2] http://www.ithome.com/html/android/86232.htm

[3] The picture. http://www.paradoxproductions.com/pics/tritwo.gif

Tuesday, May 20, 2014

Research vs. Learning



A PhD student has two roles. One is a research assistant that strives to produce good research papers. Another is a learner that keeps improving oneself. Actually, even after obtaining PhD, a researcher still need to learn, and might need to do so for one's lifetime. Research result is explicit while learning result is implicit. That why some students and faculties (e.g. [1]) ignore the later role. However, I would like to use a simple model to show that such ignorance leads to inefficiency.

Let's consider making research progress as a random sampling from a normal distribution, which represents the capability of the PhD student. And higher value of the sample corresponds to higher quality of the research. No matter how many samples are drawn, the average of the samples is the mean of the distribution. And only a few of them could have high value.

Now, student A and B both start with a normal distribution of mean = 1 and variance = 3. And we further assume that sample x > 5 means a very good paper and x < 0 means a failed research project. The figure is:

Now, A decide to solely focus on research. So A keep sampling from this distribution. However, the probability to produce a good pare is only 1.0%. So 100 trials lead to one good paper. And the chance of failure is 28.2%, so more than 1/4 of the trials result in failure.

On the other hand, B focus on learning, by which B pushed the mean of the distribution to 3:

For B, the probability of generating a good paper is 12.4%, which is better than A's in an order of magnitude. Moreover, the probability of failure now drops to 2.8%.

I don't want to draw any conclusion because it is just a very rough model. However, I think you can see the point. And I am also not saying that a PhD student should only focus on learning without any research responsibility. I personally think a PhD student should definitely spend more time on research than on learning. And doing research is actually another very important way of learning, that's why I put the Taiji graph in the beginning of this article, because they boost each other. The point I want to make is that during the journey, sometime there will be a stagnant period during research and we might feel sad. However, we should smile because as long as we keep learning and keep pushing the mean of our normal distribution to the right, things will be fine:)


References:

[1] The picture. http://www.acuherb.us/image/taiji01.png

[2] http://blog.liyiwei.org/?p=1429

Monday, May 19, 2014

Learn the Upstream

v0.1

It seems that learning the upstream of your research field would be very helpful for your research. This claim assumes that a field has a upstream field, or all fields are constructed in a hierarchical structure. I guess most people would agree with this. For example, the upstream field of Computer Science is mainly Mathematics. Mathematics provide language, tools and theorems to build the foundation of Computer Science, and many people (e.g. [1]) believe that the prerequisite of a good Computer Scientist is a solid knowledge of Mathematics. This article will expand this point to other fields by presenting several examples.

Most sub fields in computer security are more or less rooted in cryptography, which is no doubt the earliest sub field in computer security and the most rigorous one. This explains that several famous security researchers such as Ross Anderson and Bruce Schneier, started their career in cryptography, and then "invaded" many other sub fields. Andrew Yao might be another example. He switched from Physics to Computer Science, and got Turing Awards.

More examples can be found. Yin Wang has been criticizing many important products in computer science, such as SQL, Unix, Go language, ... While the validity of his criticisms are always in debate, I think they do have some value. And I further realize that it is because Yin Wang is from the programming language field, which is more or less the upstream of many other computer science subfields, such as database and OS. My roommate also serves an interesting example. Once a Math major undergraduate student, he switched to the field of Deep Learning now. Compared with researchers in CS background, his knowledge in Math helped him understand the problem deeper.

People always talk about jumping out of the box is the way towards creativity. Well I guess learning the upstream is the way towards the outside of the box. Isn't it?




References:

[1] The picture. http://www.nolandalla.com/wp-content/uploads/2014/02/salmon.jpg

[2] How to do Research At the MIT AI Lab. David Chapman. 1988

Saturday, April 19, 2014

Technological Niche

v0.1


The Facebook purchase of Oculus VR  in 2 billion dollar is still quite shocking. Particularly, it is because the company is only 2-year-old and the owner is only 21. And for people who are more familiar with technologies, there are two more reasons to be shocked. One, the legend graphics programmer, John Carmack, joined the company. And two, Virtual Reality (VR), a once hot technology, was considered a mistake and pretty much dead.

I still remember when I was an intern in Microsoft Research Asia, 2011, I attended a talk given by Professor James A. Landay. During the Q&A session, a student asked Dr. Landay about the prospect of VR. Dr. Landay said something like "many people believe VR is a mistake, and augment reality is the right way". Indeed, in the first fad of VR back in late 1990s and early 2000s, companies spent a lot of money for VR but the return was frustrating [1]. However, the founder of Oculus VR just loves VR all the time and collected most of available VR equipment. This love gives him motivation and this collection gives him inspiration. I guess most of us could regret for what unique interests we have abandoned in order to cater to others, and what large amount of money we have lost...

Let's wipe our tears and move on. Despite the lesson of perseverance, we could also find some other interesting things. For me, I would say a unique thing will never be a mistake and every unique thing will have a position and purpose in the world. And this is why I am able to write this article and you are able to read it. We know that human are mammals. But in 65 million years ago, the lord of the earth, dinosaurs, might not know mammals, which at that time, were just some tiny creatures eating bugs to survive. An alien who visited earth at that point could say that mammals are "a mistake" compared with dinosaurs. However, we all know the story afterwards, dinosaurs got wiped out in a disaster and mammals control the world. We could say that mammals win this time, but the ultimate winner is Nature who keeps the diversity of the ecosystem to combat disasters.

Same principle applies for technology. VR might not be the mainstream technology in the past 20 years and possibly not the mainstream technology in the next 20 years, as Oculus VR might fail. However, VR does have the value to exist, because it might play a critical role in the future. Every unique technology fits a technological niche, and is prepared for its day.



References:

[1] http://time.com/39577/facebook-oculus-vr-inside-story/

[2] The picture. http://media.pcgamer.com/files/2013/04/Eve-Oculus-RIft.jpg

Sunday, April 13, 2014

Resume an Old Project

v0.2


It is common for us to resume an old project (research project, programming project, etc.) due to interruptions (e.g. a vacation) or multitasking. Particularly, multitasking is necessary for graduate students and even more important for professors. The challenge is that we usually forget about the details of that project, or even forget about the motivation of doing that project. Thus, we will find quite a lot difficulties to regain what you once know. How to reduce the difficulty of resuming an old project is what we will discuss here today.

First, I found it is useful to have a warm-up period in the beginning. To recall the progress of an old research project, we could first warm up by reading a related paper. If you want to continue a suspended coding project, you can first warm up by adding some comments to the source code and refactor the code a little bit. In general, try to be slow in the beginning and do not hurry. Maybe after the warm-up and a good sleep, your dormant memories of that project start to revive and you will have more confidence to continue.

Another important thing is to write note during a project. Write down as much detail as possible as they might save a lot of time when we want to pick up where we left off. I found a simple daily note is especially useful.

Finally, before suspending a project, we need to consider how ourselves or somebody else can pick that up in the future. One tip is to make our work as automatic as possible, so the future person can quickly run it. For example, creating a script file for experiment or data analysis tasks will enable a future person to run what we've done in one command. It gives that person an instant feeling of accomplishment and it is much easier than first study a lot of stuff and then run.

References:
[1] The picture. http://blog.viddler.com/wp-content/uploads/2013/10/project-manager.jpg

Monday, April 7, 2014

Writing Like Compiling

v0.1


There has been suggestion to write programs like writing articles [2].  But I am thinking the opposite direction: could we write articles like writing programs? There is an interesting article that makes such analogy for novels [3]. While as a grinding PhD student, I am more interested in applying it to academic paper writing. And in this article, I would like to discuss the connection between writing an academic paper and compiling a program.

When a programmer has written some source code, it needs to be compiled to machine-understandable format, by a program called compiler. This is similar to writing a paper, in which you try to translate the thoughts in your mind to a form that is understandable to others. When compiling a program, there are typically many passes to process the source code and transform them step by step towards the final form. Each pass usually focuses on a specific task. Such architecture simplifies the design of compiler and enables extensions in the future. When writing a paper, we could do the same thing by first write an awful version, then improve it through multiple passes. In each pass, we focus on one goal. Below is a simple example:

  1. Make sure that the paper does not miss any important information.
  2. Make sure that the story line of the whole paper makes sense.
  3. Make sure that the core concepts are correctly defined.
  4. Make sure that terminologies are consistent and sentences are correct.
  5. Improve line by line and make the paper readable (might have a lot of redundant information)
  6. Revise the paper by removing redundant information.
  7. ...
This method definitely cannot guarantee a good paper. After all, the quality of the paper is determined by the quality of the research. However, it can at least reduce the anxiety of writers. When looking at an aghast draft, they won't feel panic and overwhelming. They can directly start from pass 1:)



References:

[1]. Picture. http://uploads3.wikipaintings.org/images/m-c-escher/drawing-hands.jpg

[2]. Literate programming. http://en.wikipedia.org/wiki/Literate_programming

[3]. 金庸笔下的良好代码风格. http://blog.sina.cn/dpool/blog/s/blog_6a55d6840101ek3y.html (In Chinese)