Monday, September 3, 2012

Relevance of Hadoop in Enterprise Application Development


RDBMS compared to hadoop/mapreduce (taken from book 'hadoop a definitive guide')

datasize - gigabytes                     v/s   petabytes
access - interactive and batch      v/s   batch
updates - read and write many times v/s write once read many times
structure - static schema              v/s   dynamic schema
integrity - high                             v/s    low
scaling - non linear                      v/s    linear

Enterprise apps are generally characterized by data volumes in gigabytes rather than petabytes, they are required to be interactive, a user makes a request and instantly reports should show it. Apps often work on transactional data (insert/update/delete), rather than write once, read always. Also enterprise apps need high data integrity.
 In an enterprise app development scenario, most applications are in control of deciding, whether the data they store will be structured or unstructured. why would they chose to store data in unstructured form? if all applications within enterprise store data in structured form, when one enterprise app needs data from another one, it will just access the structured data in that app, thru formal interfaces like direct DB access, soap, rest or a variety of other EAI options, with support for transactions and other sharing semantics.


Enterprise apps may contain data in the form of uploaded documents, etc. this can represent unstructured data, using tools like SOLR and lucene, this unstructured data can be indexed and stored off and searched relatively easily.

Within an enterprise it would not make much sense to gather stats from logs etc, unless its for purposes like finding app features usage through click model analysis. but there too, batch processes can analyse logs periodically and  summary tables can very well store the data, instead of raw logs being retained for years and years and then running hadoop based computations on the raw data

Most hadoop use cases seem to exhibit following characteristics:

  • the data under analysis seems to be unstructured and not under control of design, so that traditional RDBMS/data warehousing/ETL alternatives to hadoop cannot be employed
  • there is a need to retain the raw data as it is, so they need a cheap, scalable and commodity hardware based distributed file system anyways
  • The data is of the order of tera bytes or more rather than giga bytes
Hadoop seems to be well suited for uses such as for a internet search engine, where-in:
  • not only is there massive data of the order of terabytes and petabytes, but more importantly, 
  • this data is produced at a very fast rate, so conventional design of a batch process, summarizing this data periodically is not feasible from economics point of view
  • the massive rate at which data is being generated is not under control or determinate
  • sources which generate the data are massive in numbers and highly distributed.
  • also data sources are unknown and cannot be intercepted. eg it would not be feasible to ask twitter to for an event every time there is a tweet from someone in the world


Most enterprise applications seem to exhibit following characteristics:

  • An enterprise application produces and consumes data in various formats
  • The data produced by the application can be controlled and forced to be in structured format to some extent, including say documents uploaded by users, can be indexed and searchable
  • The rate at which data is produced is not very high like internet scale and also is not uncontrolled or indeterminate
  • data sources are well known and often can be intercepted. eg. you can ask another enterprise app to send out an async message, or poll its database or use some other notification mechanism at the source where data is generated
  • The data consumed by enterprise application is mostly from other enterprise apps, through well established EAI and notifications mechanisms
  • If an enterprise needs huge data from external sources for analytics purpose, there can be ready-made tools, COTS applications which can mine, public domain or external data and give summaries to the enterprise application. The enterprise application should not have to get that external data onboard and write their own hadoop jobs to process that data. Hadoop based analysis tools should ideally find their way into integration with enterprise applications.
  • for example if I were an automobile / finance company with a robust IT applications requirement, to get marketing data I would not think of asking my IT to use hadoop to get that data, instead, i would buy that data, or atleast use tools for getting the data, since that work is not specific to my enterprise. contrast this to having Soap or REST or EAI expertize in my IT, so that all apps in my enterprise can integrate.
Disclaimer:
I have not yet worked on any hadoop based application nor can I claim extensive experience in big data.
Still I thought putting across my thoughts based on experience in enterprise app development, might bring up questions, related to hadoop relevance in enterprise, that are also bothering minds of other enterprise architects


Note:

structured data is data that is organized into entities that have a predefined format like XML docs or database tables. this is realm of rdbms. semi structured data like spread sheets and unstructured data that does not have any internal structure like plain text or image data. map reduce works well on data that is semistructured or unstructured.

Some hadoop use cases mentioned in the book:

last.fm

  • using user generated track listening data to produce diff charts like weekly charts for top tracks per country and per user. users listen to tracks using last.fm own client or one of hundreds of third party clients apps.


facebook.com

  • producing daily and hourly summaries over large amounts of data
    • reports based on these summaries are used to drive engg and non-engg team product decisions. these summaries include reports on growth of users, page views, average time spent on site by users
    • providing performance numbers about ad campaigns that are run on facebook
    • backend processing of site features like "likes" on people, applications etc
  • running ad-hoc jobs over historical data, for analysis for product team
  • as de-facto long term archival store for log datasets
  • to lookup log events by specific attributes, to maintain site integrity and protect users against spam bots
  • complementing existing data warehousing infrastructure, by storing unstructured data


Rack Space

  • Log processing
    • hadoop is used to process logs that are generated by user interactions and end result is in lucene indexes that custmer support can query
  • to improve scalability of rdbms we resort to sharding which causes us to lose analysis of entire data. instead of sharding decided to use hadoop. scaling is linear and processing of raw data can be done parallely, using same algos for small, large or extremely large datasets



hypothetical use cases

  • advertiser insights and performance
    • advertizers have to be provided standard aggregated stats about their ads
  • ad hoc analysis and product / features feedback
  • data analysis for 
    • websites
    • bio informatics
    • oil etc explorations companies





Monday, June 18, 2012

Challenges in using HTML5 based javascript frameworks

Html5 by itself, is not a framework, so we end up using some HTML5 based frameworks. Lots of times these frameworks are also javascript based. But this has its own challenges.


Cross browser compliance is a moving target

A good "tricky" part of any javascript framework goes into trying to attain cross browser compliance. this is such a tricky deal, that javascript frameworks suffer "major" changes in moving from versions x.y.1 to x.y.2. Given the dynamic nature of javascript, application teams are wary of upgrading to newer versions of framework and some prefer to "customize" the framework. Soon, 50% of application team resources turn into "javascript/framework specialist" whose sole job is to "maintain/customize" the framework. This drains application resources and makes framework brittle while the cross browser target is still moving further and further away ;-)

Performance of ui in the browser

Due to complicated DOMs, javascript loading & execution and indiscriminate use of ajax ( I once saw a page trying to use 6 js components, each issuing its own ajax request), UI performance in browser will be major concern. Trying to build "slick" UIs, developers often incur 2-3 times server side times, in browser itself.

Javascript challenges

Javascript being a dynamic language, all issues and problems manifest at runtime, no compiler to help here.
it is oh-so-easy to override behaviour in existing frameworks and there are so many ways to achieve same result. Like all dynamic languages, power without responsibility, on part of developers can lead to code mess and undue complexity in javascript.

The regression testing nightmare

UI are enough of a nightmare for regression testing, when u add a "flexible" dynamic language like javascript to the mix, things really turn nasty.

"Single Source Componentization" difficult to achieve

Components tentacles of such frameworks are spread out over html markup, css, framework javascript and custom javascript, there is no single source UI component (in a traditional sense)


Javascript a necessary evil

.... But everything said and done, javascript frameworks are a necessary evil for web apps that are characterized by the following:
UI requirements are 'content-centric' meaning, UI is dictated by end users and will be changed and customized. Strictly component based frameworks will fragment and provide little reuse under such requests for high level customizations almost akin to a web-site UI


At the other end of the spectrum are UIs in which end users dont have much of a say, eg. operations or monitoring UIs, where arguably, the requests for customizations will not be that high and also, UI design can be dictated and fixed to reasonable extent by fewer stakeholders. For such UIs, strictly component based frameworks (as opposed to content-centric framework) might be a good match.

Types of java script based frameworks

Among the java script frameworks also there are frameworks like 

  • jquery UI which decorate existing html markup and use javascript over it to provide UI components
  • Also there is an emerging class of pure javascript frameworks which lend to componentization and MVC, like Sencha Ext.


Alternatives


Frameworks like GWT and Vaadin(call it GWT on steroids) are also viable alternatives, if you dont need too many UI customizations and thereby dont really need, control over html, css and javascript

The server side interfaces

Choose a UI framework which plugs into your server side using REST, this will ensure that the server side need not be changed at all, when you build mobile clients as alternative UIs. Additionally security considerations will be required when you think of exposing your servers over internet, so they can be accessed via your mobile clients.


Tuesday, November 15, 2011

Quick Web Security for Spring MVC based POCs

While developing POCs, often quite late we realize we have not paid attention to basic web app security features like
  • a login page or basic http auth
  • some way to specifiy mutiple users, each user having mutiple roles
  • role based access for some screens / URls in the web app
  • logout url
  • automatic redirection of non-authenticated user access to the login page 
Some of these features though quite trivial are required in the most bare of POCs, and spending development effort on this is many times not high priority
Hence the need to quicky provide about web app security features without writing a single line of code only through some basic spring security configuration.
Details as below:
1. Download spring security and put the jars in the build and runtime classpath
2. in web.xml add the following filter and filter-mapping entry
    <filter>
        <filter-name>springSecurityFilterChain</filter-name>
        <filter-class>org.springframework.web.filter.DelegatingFilterProxy</filter-class>
    </filter>
   
    <filter-mapping>
        <filter-name>springSecurityFilterChain</filter-name>
        <url-pattern>/*</url-pattern>
    </filter-mapping>
3. in web.xml, in the base spring application context, load a spring config file like security-beans.xml with following contents
     <context-param>
        <param-name>contextConfigLocation</param-name>
        <param-value>WEB-INF/existing-spring-contexts.xml,WEB-INF/security-beans.xml</param-value>
    </context-param>

contents of security-beans.xml
    <http auto-config="true" use-expressions="true">
        <intercept-url pattern="/mvc/admin/**" access="hasRole('ROLE_ADMIN')" />
        <intercept-url pattern="/mvc/general/**" access="hasRole('ROLE_SPITTER')" />
        <intercept-url pattern="/**" access="isFullyAuthenticated()" />
    </http>

    <user-service id="userService">
        <user name="all" password="all" authorities="ROLE_REGULAR_USER,ROLE_ADMIN" />
        <user name="gan" password="gan" authorities="ROLE_REGULAR_USER" />
        <user name="admin" password="admin" authorities="ROLE_ADMIN" />
    </user-service>


    <authentication-manager>
        <authentication-provider user-service-ref="userService" />
    </authentication-manager>

Explanation

Here the auto-config=true gives us a ready-made(stock) login page, which can be overriden with our own custom login page
login url: http://localhost:8081/MyWebAppContext/j_spring_security_login
logout url: http://localhost:8081/MyWebAppContext/j_spring_security_logout
userService bean allows us to specify sample userids, and their roles, accessible throughout our application through standard j2ee apis and also spring security tags on jsps
<security:authentication property="principal.username" />
<security:authorize access="hasRole('ROLE_ADMIN')">
<h2>Admin Area keep Off!</h2>
</security:authorize>

Through entry like <intercept-url pattern="/mvc/admin/**" access="hasRole('ROLE_ADMIN')" />
we can very welll control access to certain URLs in the web application for specific users and roles

Summary

Thus without writing a single line of code using spring security config we can impart quick web security to our POCs
Please refer spring security documentation for further details
 http://static.springsource.org/spring-security/site/docs/3.1.x/reference/springsecurity-single.html

Rest Integration autobinding javascript objects with controller java objects

Often when developing REST services, the server rest controller implementations work on java objects and the REST clients usually javascript ajax calls work on javascript AJAX objects. Would'nt it be great to acheive automatic binding between the javascript and java objects, instead of writing explicit code to map to and from javascript and java objects. Below is a small sample giving example of just this, using Jersey REST controllers on server side and jquery ajax request on client side. Purposely chosen a POST request scenario instead of easier GET.

Under the hoods, on server side Jersey and Jackson provide the auto binding from incoming json object to java object

 REST controller with config

    @POST
    @Path("/directUpdate")
    @Produces(MediaType.TEXT_HTML)
    @Consumes(MediaType.APPLICATION_JSON)
    public void updatePartyInfo(Party input) throws IOException{
       
        System.out.println("input param party object"+input);

        Party party = service.getByPk(input.getPartyId());
        party.setName(input.getName());
        party.setAddress(input.getAddress());
        party.setTelephone(input.getTelephone());
        party.setDescription(input.getDescription());
        service.update(party);
       

        //redirecting to subsequent page
        servletResponse.sendRedirect("http://hostname:portnum/jqueryrest/partylist.html");

    }





jQuery Ajax request on client side

    function updateParty(){
        inputData = new Object();
        inputData.partyId = $("#partyId").val();
        inputData.name = $("#name").val();
        inputData.address = $("#address").val();
        inputData.telephone = $("#telephone").val();
        inputData.description = $("#description").val();
      
        jQuery.ajax ({
            url: 'http://hostname:portnum/jqueryrest/restful/party/directUpdate',
            type: "POST",
            processData: "false",
            data: JSON.stringify(inputData),
            dataType: "json",
            contentType: "application/json",
            success: function(){
                alert('Party update success');
            }
        });

    }