Just started learning about "Hackerspace".
Without really knowing what I was doing, I fell into the 757 Labs Hackerspace.
The 757 Python Users' Group, specifically.
What a great idea. Bright people. Interested in the same area of technology.
It's like hanging around with sailors at a marina.
Rants on the daily grind of building software. This has been moved to https://slott56.github.io. Fix your bookmarks.
Moved
Moved. See https://slott56.github.io. All new content goes to the new site. This is a legacy, and will likely be dropped five years after the last post in Jan 2023.
Tuesday, June 21, 2011
Thursday, June 9, 2011
An Object-Lesson in How to Stifle Innovation
Read this: How Ma Bell Shelved the Future for 60 Years.
How many great IT projects are rejected because of this kind of delusional paranoia?
AT&T firmly believed that the answering machine, and its magnetic tapes, would lead the public to abandon the telephone.How many good ideas are set aside by managers who simply don't have a clue what users actually want?
How many great IT projects are rejected because of this kind of delusional paranoia?
Tuesday, June 7, 2011
Multithreading -- Fear, Uncertainty and Doubt
Read this: "How to explain why multi-threading is difficult".
We need to talk. This is not that difficult.
Multi-threading is only difficult if you do it badly. There are an almost infinite number of ways to do it badly. Many magazines and bloggers have decided that the multithreading hurdle is the Next Big Thing (NBT™). We need new, fancy, expensive language and library support for this and we need it right now.
Parallel Computing is the secret to following Moore's Law. All those extra cores will go unused if we can't write multithreaded apps. And we can't write multi-threaded apps because—well—there are lots of reasons, split between ignorance and arrogance. All of which can be solved by throwing money after tools. Right?
Arrogance
One thing that makes multi-threaded applications error-prone is simple arrogance. There are lots and lots of race conditions that can arise. And folks aren't trained to think about how simple it is to have a sequence of instructions interrupted at just the wrong spot. Any sequence of "read, work, update" operations will have threads doing reads (in any order), threads doing the work (in any order) and then doing the updates in the worst possible order.
Compound "read, work, update" sequences need locks. And the locations of the locks can be obscure because we rarely think twice about reading a variable. Setting a variable is a little less confusing. Because we don't think much about reads, we fail to see the consequences of moving the read of a variable around as part of an optimization effort.
Ignorance
The best kind of lock is not a mutex or a semaphore. It surely isn't an RDBMS (but God knows, numerous organizations have used an RDBMS as a large, slow, complex and expensive message queue.)
The best kind of lock seems to be a message queue. The various concurrent elements can simply dequeue pieces of data, do their tasks and enqueue the results. It's really elegant. It has many, simple, uncoupled pieces. It can be scaled by increasing the number of threads sharing a queue.
A queue (read with an official "get") means that the reads aren't casually ignored and moved around during optimization. Further, the creation of a complex object can be done by one thread which gets pieces of data from a queue shared by multiple writers. No locking on the complex object.
Using message queues means that there's no weird race condition when getting data to start doing useful work; a get is atomic and guaranteed to have that property. Each thread gets an thread-local, thread-safe object. There's no weird race condition when passing a result on to the next step in a pipeline. It's dropped into the queue, where it's available to another thread.
Dining Philosophers
The Dining Philosophers Code Kata has a queue-based solution that's pretty cool.
A queue of Forks can be shared by the various Philosopher threads. Each Philosopher must get two Fork resources from the queue, eat, philosophize and then enqueue the two Forks again. It's quite short, easy to write and easy to demonstrate that it must work.
Perhaps the hardest thing is designing the Dining Room (also know as the Waiter, Conductor or Footman) that only allows four of the five philosophers to dine concurrently. To do this, a departing Philosopher must enqueue themselves into a "done eating" queue so that the next waiting Philosopher can be seated.
A queue-based solution is delightfully simple. 200 or so lines of code including docstrings comments so that the documentation looked nice, too.
Additional Constraints
The simplest solution uses a single queue of anonymous Forks. A common constraint is to insist that each Philosopher use only the two adjacent forks. Philosopher p can use forks (p+1 mod 5) and (p-1 mod 5).
This is pleasant to implement. The Philosopher simply dequeues a fork, checks the position, and re-enqueues it if it's a wrong fork.
FUD Factor
I think that the publicity around parallel programming and multithreaded applications is designed to create Fear, Uncertainty and Doubt (FUD™).
- Too many questions on StackOverflow seem to indicate that a slow program might magically get faster if somehow threads where involved. For programs that involve scanning the entire hard drive or downloading Wikipedia or doing a giant SQL query, the number of threads has little relevance to the real work involved. These programs are I/O bound; since threads must share the I/O resources of the containing process, multi-threading won't help.
- Too many questions on StackOverflow seem to have simple message queue solutions. But folks seem to start out using inappropriate technology. Just learn how to use a message queue. Move on.
- Too many vendors of tools (or languages) are pandering to (or creating) the FUD factor. If programmers are made suitably fearful, uncertain or doubtful, they'll lobby for spending lots of money for a language or package that "solves" the problem.
Sigh. The answer isn't software tools, it's design. Break the problem down into independent parallel tasks and feed them from message queues. Collect the results in message queues.
Some Codeclass Philosopher( threading.Thread ):
"""A Philosopher. When invited to dine, they will
cycle through their standard dining loop.
- Acquire two forks from the fork Queue
- Eat for a random interval
- Release the two forks
- Philosophize for a random interval
When done, they will enqueue themselves with
the "footman" to indicate that they are leaving.
"""
def __init__( self, name, cycles=None ):
"""Create this philosopher.
:param name: the number of this philosopher.
This is used by a subclass to find the correct fork.
:param cycles: the number of cycles they will eat.
If unspecified, it's a random number, u, 4 <= u < 7
"""
super( Philosopher, self ).__init__()
self.name= name
self.cycles= cycles if cycles is not None else random.randrange(4,7)
self.log= logging.getLogger( "{0}.{1}".format(self.__class__.__name__, name) )
self.log.info( "cycles={0:d}".format( self.cycles ) )
self.forks= None
self.leaving= None
def enter( self, forks, leaving ):
"""Enter the dining room. This must be done before the
thread can be started.
:param forks: The queue of available forks
:param leaving: A queue to notify the footman that they are
done.
"""
self.forks= forks
self.leaving= leaving
def dine( self ):
"""The standard dining cycle:
acquire forks, eat, release forks, philosophize.
"""
for cycle in range(self.cycles):
f1= self.acquire_fork()
f2= self.acquire_fork()
self.eat()
self.release_fork( f1 )
self.release_fork( f2 )
self.philosophize()
self.leaving.put( self )
def eat( self ):
"""Eating task."""
self.log.info( "Eating" )
time.sleep( random.random() )
def philosophize( self ):
"""Philosophizing task."""
self.log.info( "Philosophizing" )
time.sleep( random.random() )
def acquire_fork( self ):
"""Acquire a fork.
:returns: The Fork acquired.
"""
fork= self.forks.get()
fork.held_by= self.name
return fork
def release_fork( self, fork ):
"""Acquire a fork.
:param fork: The Fork to release.
"""
fork.held_by= None
self.forks.put( fork )
def run( self ):
"""Interface to Thread. After the Philosopher
has entered the dining room, they may engage
in the main dining cycle.
"""
assert self.forks and self.leaving
self.dine() The point is to have the dine method be a direct expression of the Philosopher's dining experience. We might want to override the acquire_fork method to permit different fork acquisition strategies.
For example, a picky philosopher may only want to use the forks adjacent to their place at the table, rather than reaching across the table for the next available Fork.
The Fork, by comparison, is boring.
The Table, however, is interesting. It includes the special "leaving" queue that's not a proper part of the problem domain, but is a part of this particular solution.
The dinner method assures that all Philosophers eat until they are finished. It also assures that four Philosophers sit at the table and when one finishes, another takes their place. Finally, it also assures that all Philosophers are done eating before the dining room is closed.
For example, a picky philosopher may only want to use the forks adjacent to their place at the table, rather than reaching across the table for the next available Fork.
The Fork, by comparison, is boring.
class Fork( object ):
"""A Fork. A Philosopher requires two of these to eat."""
def __init__( self, name ):
"""Create the Fork.
:param name: The number of this fork. This may
be used by a Philosopher looking for the correct Fork.
"""
self.name= name
self.holder= None
self.log= logging.getLogger( "{0}.{1}".format(self.__class__.__name__, name) )
@property
def held_by( self ):
"""The Philosopher currently holding this Fork."""
return self.holder
@held_by.setter
def held_by( self, philosopher ):
if philosopher:
self.log.info( "Acquired by {0}".format( philosopher ) )
else:
self.log.info( "Released by {0}".format( self.holder ) )
self.holder= philosopher
The Table, however, is interesting. It includes the special "leaving" queue that's not a proper part of the problem domain, but is a part of this particular solution.
class Table( object ):
"""The dining Table. This uses a queue of Philosophers
waiting to dine and a queue of forks.
This sets Philosophers, allows them to dine and then
cleans up after each one is finished dining.
To prevent deadlock, there's a limit on the number
of concurrent Philosophers allowed to dine.
"""
def __init__( self, philosophers, forks, limit=4 ):
"""Create the Table.
:param philosophers: The queue of Philosophers waiting to dine.
:param forks: The queue of available Forks.
:param limit: A limit on the number of concurrently dining Philosophers.
"""
self.philosophers= philosophers
self.forks= forks
self.limit= limit
self.leaving= Queue.Queue()
self.log= logging.getLogger( "table" )
def dinner( self ):
"""The essential dinner cycle:
admit philosophers (to the stated limit);
as philosophers finish dining, remove them and admit more;
when the dining queue is empty, simply clean up.
"""
self.at_table= self.limit
while not self.philosophers.empty():
while self.at_table != 0:
p= self.philosophers.get()
self.seat( p )
# Must do a Queue.get() to wait for a resource
p= self.leaving.get()
self.excuse( p )
assert self.philosophers.empty()
while self.at_table != self.limit:
p= self.leaving.get()
self.excuse( p )
assert self.at_table == self.limit
def seat( self, philosopher ):
"""Seat a philosopher. This increments the count
of currently-eating Philosophers.
:param philosopher: The Philosopher to be seated.
"""
self.log.info( "Seating {0}".format(philosopher.name) )
philosopher.enter( self.forks, self.leaving)
philosopher.start()
self.at_table -= 1 # Consume a seat
def excuse( self, philosopher ):
"""Excuse a philosopher. This decrements the count
of currently-eating Philosophers.
:param philosopher: The Philosopher to be excused.
"""
philosopher.join() # Cleanup the thread
self.log.info( "Excusing {0}".format(philosopher.name) )
self.at_table += 1 # Release a seat
The dinner method assures that all Philosophers eat until they are finished. It also assures that four Philosophers sit at the table and when one finishes, another takes their place. Finally, it also assures that all Philosophers are done eating before the dining room is closed.
Friday, June 3, 2011
Changed the Page Template
The "default" template I chose was too narrow for presenting code samples. Changed it.
Thursday, May 26, 2011
Code Kata : "Simple" Database Design
Here's a pretty simple set of use cases for a code-kata database application.
This is largely transactional, not analytical.
It's a simple inventory of ingredients, recipes and locations.
Context
- 42' sailboat.
- Lots of places to keep stuff. Lots.
Stuff gets lots or misplaced. It's helpful to marry recipes with ingredients to use up the last of something before it goes bad and stinks up the boat.
Actor is essentially the cook.
Use Cases
- Perishables to be eaten soon?
- Shopping list for specific recipes.
- Where did I put that?
Model
- Ingredient. A generic description: "lime", "coconut". Not too much more is needed. A "food safety" notation (refrigeration required, etc.) is a helpful attribute. Maybe a "food group" or other nutrition information.
- Location. A text description of where things can be stored. This shouldn't have too many attributes, because boats aren't big grids. Phrases like "port saloon upper cabinet", or "galley outer cooler" make sense to folks who live on the boat.
- On Hand. This is simply ingredient, location and a measurement of some kind. Example: 3 limes in the starboard galley center cooler. There's a lot of magic around units and unit conversion that can be fun. But that strays outside the database domain.
- Recipe. Example: "One of sour, two of sweet, three of strong, and four of weak.", lime, simple syrup, rum, water. Plain text using a lightweight markup is what's required here. Along with a many-to-many relationship with ingredients. This is not carefully defined above because it should be done as a "more advanced" exercise.
I think this has the right amount of complexity and isn't very abstract. Since the use cases are pretty obvious to anyone who's cooked or been to a grocery store, use case details aren't essential.
Wednesday, May 25, 2011
Meetup Tonight
Tonight (May 25th). Red Dog. Colley Ave. Ghent. I'll be wearing my Stack Overflow shirt. I'll be there about 7. I know that at least one other person won't be there until 8.
The Meetup link.
I like this meetup idea a lot. Probably because the WFH life-style is a little isolating.
There's the small "Hampton Stack Overflow Community". We have a common interest in Stack Overflow.
Also, there's the 757 Python Users Group. We have a common interest in Python. I've decided to become the "official" organizer for this. I'm going to join the 757 Labs Hackerspace, also.
Tuesday, May 24, 2011
Agility and following a "Strictly Agile" approach
I've seen some discussion on Stack Overflow that is best characterized by the question: "What is Strictly Agile?", or "What's the Official Agile Approach?".
Someone shared this with me recently: "Process kills developer passion".
I have also heard some great complaints about organizations that claim "Agile" and actually do nothing of the kind. In some cases it's not a "crunchy agile shell" around a waterfall process; it's a simple lie. Nothing about the process is Agile except a manager insisting that all the status reporting, planning and unprioritized lists of random requirements are Agile.
Finally, I got this weird suggestion: "consider writing a blog about how to test if you are agile or not". It's weird because testing for Agile is like testing for breathing; it's like testing for flammability.
The Agile Test
Testing if your project is Agile can be done two ways.
- Practical. Make a change to the project. Any change. Requirements, architecture, due dates, staff, anything. Does it derail? If so, it wasn't very Agile, was it?
- Theoretical. Reread the Agile Manifesto. Make a score card that evaluates the project on each of the eight basic criteria in the Agile manifesto. Convene all the project stakeholders. Conduct careful surveys and have structured walkthroughs to determine the degree of Agility surrounding each person, deliverable, collaborative relationship and issue.
An important point is that Agile is not absolute. Some practices are more Agile than others. There's no "strictly" Agile. There are ways to make a project more Agile; that is, it can effectively cope with change. There are ways to make a project less Agile; that is, change causes problems and can derail the project completely.
The canonical example is a missing, misstated or contradictory requirement that gets uncovered after coding and during user acceptance test. Clearly, that feature has been built and is absolutely wrong. What happens next?
Agile? The product can be released with with the broken feature relegated to the next release. A hack is put in to remove the buttons or menu items or links until they work.
Not Agile? Everyone works around the clock to make that feature work no matter what. Paraphrasing Admiral Farragut: "Technical debt be damned. Development must proceedfull speed ahead." All of this irrespective of the relative value of what's being developed. Schedule comes first; features second.
How Much Process?
The "Process Kills..." blog entry repeats observation that a lot of carefully-defined process isn't really all that helpful. It identifies a cause ("process kills passion") that's can be true, but it's largely irrelevant. Process is—essentially—work that's not focused on delivering anything of real value. Complex processes are "meta" work; it's work focused on IT internals; it's work that creates no value for the users of the software; work that replaces the more valuable elements of the Agile Manifesto.
One can argue that processes, documentation, contracts and plans "assure" success or demonstrate some level of quality. To an extent all the process and meta-work creates trust that—eventually—the resulting software product will solve the original problem.
The mistake is that non-Agile methods use a series of surrogates—processes, documentation, contracts and plans—instead of actual software. The point of Agile methods it to release software early and often and avoid using surrogates.
Key Points of Agile
Here are the key points of the Agile Manifesto.
- Individuals and interactions over processes and tools. A more Agile project will use the best people and encourage them to talk amongst themselves. A less Agile project will write a lot of things (which folks don't have time or reward for reading.) There will be misunderstandings, leading to large, boring meetings where someone reads powerpoint slides to other folks to try and clear up misunderstandings.
- Working software over comprehensive documentation. A more Agile project uses frequent release cycles of incremental software. A less Agile project attempts to gather all requirements, do all design and then try to do all the coding even though the requirements have already been found to be less than crystal clear.
- Customer collaboration over contract negotiation. A more Agile project uses constant contact with customer and product owner to refine and prioritize the requirements. A less Agile project uses a complex change control process to notify everyone of a requirements change, which leads to design and code changes, and has cost and schedule impact that must be carefully planned and documented.
- Responding to change over following a plan. A more Agile project uses incremental releases, conversation and a modicum of discipline to build things of value. Just because someone thought it should be included in the requirements doesn't mean the feature is really required.
The "Process Kills Passion?" Question
There Process Kills Passion blog lists a bunch of things that—it appears—some folks find burdensome:
- Doing full TDD, writing your tests before you wrote any implementing code.
- Requiring some arbitrary percentage of code coverage before check-in.
- Having full code reviews on all check-ins.
- Using tools like Coverity to generate code complexity numbers and requiring developers to refactor code that has too high a complexity rating.
- Generating headlines, stories and tasks.
- Grooming stories before each sprint.
- Sitting through planning sessions.
- Tracking your time to generate burn-down charts for management.
This list has three different collections of practices.
- Good. TDD, code reviews, generating headlines, stories and tasks, grooming stories before each sprint and doing some planning for each sprint are all simply good ideas. They must be done. "Pure Coding" is not a good way to invest time. Planning and then coding is much smarter, no matter how boring planning appears.
- Difficult. Test code coverage can be helpful, but can also devolve to empty numerosity. 20% more coverage doesn't not mean 20% fewer bugs. Nor does it mean 20% less chance of uncovering a bug at run time. Code complexity ratings are also fussy because they don't have a direct correlation with much. They must be done and used to prioritize work that will reduce technical debt. But mindless thresholds are for cowards who don't want to mediate deep technical discussions.
- Silly. Creating burn-down charts for management shouldn't be necessary. Everyone must read and understand the backlog. Everyone should build the summary charts they want from the backlog. The product owner or even the eventual customer should do this on their own. They must be given a profound level of ownership of the features and the process for creating software.
I don't agree that process kills passion. I think there's a fine line between playing with software development and building software of value. I think that valuable software requires some discipline and requires executing a few burdensome tasks (like TDD) that create real value. Assuring 80% or 100% code coverage doesn't always create real value. Spending time keeping the backlog precise and complete is good; spending time making pictures is less good.
Subscribe to:
Posts (Atom)