Main Article Content
Abstract
This paper focuses on identifying the unusual sizes of images in Big Data. As the term Big Data itself refers, it is already Big and it still will go bigger and bigger through many channels like audio, video, web sites, advertisements, social media and many. Companies are investing huge amount of money for storing the raw data assuming that all the data will be useful and it is of high importance. In the centralized data servers, multiple copies of the same document / file which are getting stored may also create unwanted junk of big data. Considering the image data, the scale invariants problem may be overcome by removing very small-scale images as it is taking much processing time and having less accuracy. This paper helps in removing the unusual small-scale images using the simple ad-hoc method. All the image features are extracted using two standard feature extraction techniques i.e., Gray Level Co-occurrence Matrix and Zernike Moments. Then, the standard deviation on those features is exactly identifying the unusual images, which can be the big-scale, or the small-scale images. The results are proving the effectiveness of the proposed method. The work is implemented on Matlab 7.0 and tested on the “Yahoo Flickr Two Dataset”.