Showing posts with label maintenance. Show all posts
Showing posts with label maintenance. Show all posts

Next generation refactoring with syntactical grep and patch from pfff

Refactoring of PHP methods is often difficult:
  • syntax errors or non-existing methods are only detected during runtime
  • wrong method calls or uninitialized variables are only detected during runtime
  • wrong order of parameters often remains undetected
  • not enough unit tests to validate the changes
  • not many resources for refactoring (time + money)
We can solve most of these issues by using a few tools for syntactic analysis.

You might have already worked with grep to search your code:


# find a function in all PHP files (and sub-directories)
grep -rin "SomeFunc(" *.php

This also finds "doSomeFunc(" as well as strings or documentation containing "SomeFunc(".
If you want to find all occurrences of SomeFunc() with exactly 2 parameters or at least 1 parameter, things get complicated.

Happily, Facebook has released a few tools to run static analysis and source-to-source transformations on PHP code. The package is named pfff and is available on GitHub.

Syntactic grep

Using sgrep from pfff, we can do syntactic searches in the code:


# find all occurrences of SomeFunc() with exactly 2 parameters
sgrep -e "SomeFunc(X,Y)" *.php

# find all occurrences of SomeFunc() with at least 1 parameter
sgrep -e "SomeFunc(X,...)" *.php

# ... is a wildcard for any number of parameters
# X, Y are wildcards for a parameter, X != Y

Syntactic patch

Using spatch, we can use syntactic searches to refactor our code:


# change parameters order from ABC to CAB
spatch -e 's/SomeFunc(A,B,C)/SomeFunc(C, A, B)/' *.php

# drop second parameter
spatch -e 's/SomeFunc(A,B,C)/SomeFunc(A, C)/' *.php

# rename SomeFunc to SomeOtherFunc
spatch -e 's/SomeFunc(...)/SomeOtherFunc(...)/' *.php

Install pfff

There are currently no binary packages, so you need to compile pfff manually on your machine. Here is a small guide to get it done with Ubuntu 12.10:


# ocaml 4.0 is currently only in Debian experimental
echo deb http://ftp.de.debian.org/debian experimental main \
>/etc/apt/sources.list.d/debian_exp.list
apt-get install git build-essential debian-archive-keyring
apt-get update
apt-get install libpcre3-dev libgtk2.0-dev binutils-gold gawk
apt-get install -t experimental ocaml camlp4 ocaml-base ocaml-nox \
ocaml-base-nox ocaml-interp ocaml-compiler-libs
cd /
git clone --depth=1 git://github.com/facebook/pfff.git
cd pfff
./configure
make depend && make && make opt
make install
Note: After compiling pfff, the binaries of sgrep and spatch can be directly copied to other systems without installing or compiling other packages.

Resources:

Applying Scrum to legacy code and maintenance tasks

There are some problems with Scrum that mainly occur when dealing with third party components or legacy systems:
  • wrong estimations (time, impact, risk, complexity)
  • bad requirements (inconsistent, incomplete, testable, conflicting, faulty)
  • development involved in operations (bug analysis, data correction, deployment, hot-fixes)
  • delays in development (bugs in legacy system, missing documentation)
  • testing (quality/quantity of test cases, un-mockable interfaces, long running offline processes, performance issues, live and test system differ)

General problems with Scrum:
  • development (gold plating, rework, misunderstandings and bugs from collective code ownership)
  • estimations (missing experience, excessive estimations for unpleasant stories)
  • technical debt vs. velocity (architecture violations save deadlines)
  • performance (stories are functional requirements, performance is normally no acceptance criteria, often only specified as "system should be fast and responsive")
  • testing (time, cost, hidden bugs, product owner focused on functionality and business value, not code quality)
  • absence, outstanding feedback (illness, vacation, conflicts, laziness)

This gives some questions:
  • Are all these problems exceptions or daily business in software development?
  • Does Scrum deal with these problems efficiently or solve them?
  • Does Scrum only work in an ideal world?

Some discussions on these topics can be found here:
From the theoretical point of view, the main issue is synchronizing all stories and developers at the end of the sprint. Stopping a sprint for a hot-fix is quite uncommon, instead management just overloads developers. Developers are forced to hold deadlines, even if estimations tend to be wrong. This can lead to more bugs, less tests/documentation, incomplete stories, more hot-fixes, more delays, developers being idle and others being overloaded. Some people skip estimations and testing to make scheduling of deadlines easier, but quality is not for free!
Doing things like "insert at least one task to each sprint to improve the process" or "insert at least one task to each sprint to avoid the code becoming legacy code" is the way to go. Same for evaluating new technologies or tools. Other adaptions for quality in Scrum are "insert a testing sprint before the rollout" or setting up teams cross-functional, which means they should not include only developers, but also testers, configuration and requirements managers.

Feature-driven development defines features without a sprint scope and therefore has less dependencies. A feature/developer can be paused to do a hot-fix without any impact on other features. But having 2 features that depend on each other still requires completeness for both.

Disadvantages of ORM

ORM has attracted a lot of attention in the last years. So let's get a bit deeper into it.

The biggest advantage of ORM is also the biggest disadvantage: queries are generated automatically
  • queries can't be optimized
  • queries select more data than needed, things get slower, more latency
    (some ORMs fetch all datasets of all relations of an object even though only 1 attribute is read)
  • compiling queries from ORM code is slow (ORM compiler written in PHP)
  • SQL is more powerful than ORM query languages
  • database abstraction forbids vendor specific optimizations

Other problems coming up with ORM
  • compiling ORM logic from phpDoc instructions or XML files is slow, but can be cached
  • ORM validates relations and field names outside the database, but can't keep relations consistent
  • ORM libraries are often used in projects without making a benchmark before
  • ORM libraries are often used because the documentation of the library says it is very fast
  • ORM libraries are often used by default without checking the project's needs
  • database abstraction is often required but changing the database never happens
  • databases are not object oriented
  • ORM violates the basic database performance principle: you get the best performance when your data is stored in the same structure it gets read

General coding problems with ORM
  • having objects instead of SQL, programmers tend to write joins directly in PHP
  • ORM code can be much longer than normal code with PHP and SQL
    (increase of complexity, error rates and maintenance efforts)
  • how to handle null values? (assign null => isset gives false)
  • people often document PHP code but not the database schemas
    (e.g. empty comments in MySQL fields and tables, docs not up-to-date)
  • new versions of ORM libraries often forbid reusing older ORM code
  • slow code is often wrapped with caching, so you always serve old data

Where can ORM be good?
  • avoid building SQL strings for simple insert, update, delete
  • using ORM with magic getters/setters in PHP
  • allow models to inherit attributes and methods from other models
  • separate models from views and controllers
  • centralize validation rules, save or delete methods to one class per entity
  • handle escaping and serialization of values automatically

Performance in numbers?
e.g. Doctrine 2, watch slide 50 and 54: Doctrine is >3 times slower than raw PHP on 20 inserts, imagine what happens with 20000 ... real numbers are much slower, see slide 47, here the authors only benchmarked flush() instead of the whole code

Coming soon: How to write a really small and fast O/R-mapper with PHP